Transformers without Normalization
Jiachen Zhu, Xinlei Chen, Kaiming He +2 authors
Dynamic Tanh (DyT) replaces normalization layers in Transformers, achieving equivalent or superior performance without hyperparameter tuning across various tasks.
Explore · 每周精选
发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。
Kristian Kuznetsov, Laida Kushnareva, Polina Druzhinina +5 authors
Sparse Autoencoders enhance interpretability in Artificial Text Detection by extracting distinctive features from LLM outputs, providing insights into differences from human-written texts.
50 篇论文 · 按点赞排序
Jiachen Zhu, Xinlei Chen, Kaiming He +2 authors
Dynamic Tanh (DyT) replaces normalization layers in Transformers, achieving equivalent or superior performance without hyperparameter tuning across various tasks.
Aleksandr Nesterov, Andrey Sakhovskiy, Ivan Sviridov +5 authors
Experiments on a new Russian-language ICD coding dataset using models like BERT, LLaMA with LoRA, and RAG show significant accuracy improvements in automated clinical coding compared to manual annotations.
Yibin Wang, Yuhang Zang, Hao Li +2 authors
A unified reward model for multimodal understanding and generation improves the assessment and preference alignment in image and video tasks through joint learning and direct preference optimization.
Samuel Cahyawijaya, Holy Lovenia, Joel Ruben Antony Moniz +89 authors
SEA-VL is an open-source project aimed at creating a large, culturally relevant dataset for Southeast Asian languages to improve AI inclusivity.
Eliahu Horwitz, Nitzan Kurer, Jonathan Kahana +2 authors
An atlas of Hugging Face's model repository provides visualizations and analysis, with methods for mapping undocumented areas based on structural priors.
Yingzhe Peng, Gongrui Zhang, Miaosen Zhang +7 authors
A two-stage framework enhances reasoning in large multimodal models by first strengthening text-only reasoning and then generalizing it to multimodal tasks.
Nicolas Boizard, Hippolyte Gisserot-Boukhlef, Duarte M. Alves +16 authors
EuroBERT, a family of multilingual encoders covering European and global languages, outperforms existing models across various tasks and supports long sequences, surpassing traditional bidirectional encoders.
Advait Gupta, NandaKiran Velaga, Dang Nguyen +1 authors
A three-stage approach, CoSTA*, combines LLMs and A* search to efficiently find cost-effective tool paths for multi-turn text-to-image editing, outperforming state-of-the-art models in cost and quality.
Marianne Arriola, Aaron Gokaslan, Justin T Chiu +5 authors
Block diffusion language models improve generation efficiency and sequence length flexibility compared to autoregressive and discrete denoising diffusion models.
Ruibin Yuan, Hanfeng Lin, Shuyue Guo +54 authors
YuE, a family of open foundation models based on LLaMA2, can generate long-form music with aligned lyrics, coherent structure, and appropriate accompaniment using innovative techniques in next-token prediction, conditioning, and pre-training.
Xun Liang, Hanyu Wang, Huayi Lai +7 authors
Sparse Expert Activation Pruning (SEAP) is a method for pruning large language models that reduces computational overhead while maintaining accuracy by identifying and retaining task-specific expert activations.
Fanqing Meng, Lingxiao Du, Zongkai Liu +11 authors
MM-Eureka extends rule-based reinforcement learning to multimodal reasoning, achieving strong capabilities in data-efficient multimodal tasks without supervised fine-tuning.
Zeyinzi Jiang, Zhen Han, Chaojie Mao +3 authors
VACE, an all-in-one framework for video creation and editing, integrates multiple tasks within a unified model using a Video Condition Unit and Context Adapter for flexible and consistent video synthesis.
Siyin Wang, Zhaoye Fei, Qinyuan Cheng +4 authors
Dual Preference Optimization (D$^2$PO) enhances embodied task planning in large vision-language models by jointly optimizing state prediction and action selection using preference learning and tree search for efficient trajectory collection.
Hengguang Zhou, Xirui Li, Ruochen Wang +3 authors
A non-SFT model replicated emergent reasoning characteristics for multimodal tasks using reinforcement learning, achieving higher accuracy than base and SFT models on CVBench.
Rongyao Fang, Chengqi Duan, Kun Wang +9 authors
A new reasoning-driven paradigm, GoT, improves text-to-image generation and editing through language understanding and reasoning chains, integrating Qwen2.5-VL and an enhanced diffusion model with a Semantic-Spatial Guidance Module.
Jinhyuk Lee, Feiyang Chen, Sahil Dua +44 authors
Gemini Embedding, utilizing Google's Gemini large language model, generates high-quality multilingual and code embeddings outperforming benchmarks across various tasks.
Yuxiao Qu, Matthew Y. R. Yang, Amrith Setlur +4 authors
The paper formalizes test-time compute optimization as a meta-reinforcement learning problem, introducing Meta Reinforcement Fine-Tuning (MRT) to enhance performance and token efficiency in large language models.
Feng Jiang, Zhiyu Lin, Fan Bu +3 authors
S2S-Arena is introduced to evaluate speech models' instruction-following abilities with paralinguistic information, revealing that cascaded ASR, LLM, and TTS outperform jointly trained models in speech2speech protocols and that generating appropriate audio with paralinguistic information remains challenging.
Simon A. Aytes, Jinheon Baek, Sung Ju Hwang
A prompting framework named Sketch-of-Thought combines cognitive reasoning paradigms to reduce token usage in large language models with minimal impact on accuracy.
Lingmin Ran, Mike Zheng Shou
A multi-stage diffusion framework, TPDiff, enhances video diffusion model efficiency by reducing full frame rate during high-entropy stages, leading to diminished training costs and improved inference efficiency.
Weijia Wu, Zeyu Zhu, Mike Zheng Shou
MovieAgent automates movie generation through a hierarchical Chain of Thought framework using multiple language models to handle narrative structuring, ensuring script fidelity, character consistency, and narrative coherence.
Junsong Chen, Shuchen Xue, Yuyang Zhao +6 authors
SANA-Sprint is an efficient diffusion model for ultra-fast text-to-image generation, utilizing hybrid distillation and integrating ControlNet for interactive real-time image generation with high-quality outputs and minimal latency.
Bowen Jin, Hansi Zeng, Zhenrui Yue +3 authors
Search-R1, an extension of DeepSeek-R1, enhances large language models by autonomously generating multiple search queries during step-by-step reasoning using reinforcement learning, improving performance in retrieval-augmented question-answering tasks.
Jiazheng Liu, Sipeng Zheng, Börje F. Karlsson +1 authors
MMDiag, a multi-turn multimodal dialogue dataset, challenges MLLMs with real-world conversational scenarios, and DiagNote, an MLLM with multimodal grounding and reasoning, outperforms existing models in these tasks.
Jiaxing Zhao, Xihan Wei, Liefeng Bo
Reinforcement Learning with Verifiable Reward applied to a multimodal large language model enhances emotion recognition, reasoning, and generalization.
Lixue Gong, Xiaoxia Hou, Fanshi Li +25 authors
Seedream 2.0, a bilingual Chinese-English image generation model, excels in text rendering and cultural nuances, employing data integration, caption balancing, and multi-phase optimizations to achieve state-of-the-art performance.
Weiyun Wang, Zhangwei Gao, Lianjie Chen +12 authors
VisualPRM enhances multimodal reasoning capabilities in MLLMs using Best-of-N strategies and a process supervision dataset, surpassing other models in benchmarks.
Hongwei Yi, Tian Ye, Shitong Shao +10 authors
MagicInfinite, a diffusion Transformer framework, provides high-fidelity portrait animations across various styles with techniques like 3D full-attention, curriculum learning, region-specific masks, and step and cfg distillation.
北京市昌平区探索星信息技术及软件开发工作室
京ICP备2026059466号