Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
Tri Dao, Albert Gu
A new framework connects state-space models and transformers, leading to a faster architecture for language modeling.
Explore · 每周精选
发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。
Lin Chen, Xilin Wei, Jinsong Li +12 authors
The ShareGPT4Video series enhances video understanding and generation using dense captions and an efficient captioning strategy by addressing temporal and spatial challenges in video annotation.
41 篇论文 · 按点赞排序
Tri Dao, Albert Gu
A new framework connects state-space models and transformers, leading to a faster architecture for language modeling.
Yubo Wang, Xueguang Ma, Ge Zhang +14 authors
MMLU-Pro extends the MMLU benchmark with more challenging reasoning questions, eliminates trivial ones, and demonstrates better stability and discriminative power for language models.
Namgyu Ho, Sangmin Bae, Taehyeon Kim +6 authors
The Block Transformer architecture enhances inference throughput by applying global-to-local modeling to autoregressive transformers, reducing inference bottlenecks through hierarchical processing and block-level self-attention.
Philip Anastassiou, Jiawei Chen, Jitong Chen +43 authors
Seed-TTS is a family of large-scale TTS models that generate high-quality speech with in-context learning, superior controllability, and a non-autoregressive variant using diffusion-based architecture that does not rely on pre-estimated phoneme durations.
Yang Sui, Yanyu Li, Anil Kag +7 authors
A weight quantization method reduces the size of diffusion-based image generation models while improving generation quality.
Hai-Long Sun, Da-Wei Zhou, Yang Li +8 authors
Parrot enhances multimodal language models with multilingual visual token alignment using textual guidance and Mixture-of-Experts, achieving state-of-the-art performance on multilingual benchmarks.
Yasin Abbasi Yadkori, Ilja Kuzborskij, András György +1 authors
Information-theoretic uncertainty quantification in large language models reliably detects epistemic uncertainty, allowing detection of hallucinations in both single- and multi-answer responses.
Junyang Wang, Haiyang Xu, Haitao Jia +6 authors
Mobile-Agent-v2, a multi-agent system with planning, decision, and reflection components, improves task completion in mobile device operations by addressing navigation challenges and handling errors.
Omar Shaikh, Michelle Lam, Joey Hejna +3 authors
Demonstration ITerated Task Optimization (DITTO) aligns language models to specific settings using very few demonstrations, outperforming few-shot prompting and supervised fine-tuning across various domains.
Ling Yang, Zhaochen Yu, Tianjun Zhang +5 authors
Buffer of Thoughts (BoT) enhances large language models with a meta-buffer of adaptive thought-templates, improving performance, efficiency, and robustness across various reasoning tasks.
Zhanhao Liang, Yuhui Yuan, Shuyang Gu +4 authors
Step-aware Preference Optimization improves alignment of text-to-image diffusion models with human preferences by independent evaluation and adjustment at each denoising step, outperforming existing methods in aesthetics and efficiency.
Zhixing Zhang, Yanyu Li, Yushu Wu +9 authors
Adversarial training transforms a multi-step diffusion video generation model into a single-step high-quality video synthesis model with reduced computational cost.
Ye Tian, Ling Yang, Haotian Yang +9 authors
VideoTetris framework uses compositional diffusion and enhanced preprocessing to generate complex text-to-video content with improved consistency and dynamic handling.
Chaoyou Fu, Yuhan Dai, Yondong Luo +17 authors
Video-MME is a comprehensive, high-quality benchmark that evaluates Multi-modal Large Language Models in video analysis, covering diverse video types, durations, modalities, and precise annotations.
Zhiheng Xi, Yiwen Ding, Wenxiang Chen +17 authors
AgentGym is a framework that facilitates the development of self-evolving LLM-based agents capable of handling diverse tasks across various environments without extensive human supervision.
Zachary Ankner, Cody Blakeney, Kartik Sreenivasan +3 authors
Perplexity-based data pruning using small language models improves performance and reduces training time for larger models across various datasets and conditions.
Jiahao Shao, Yuanbo Yang, Hongyu Zhou +4 authors
A conditional generation approach using a video diffusion model optimizes spatial and temporal layers for improved depth estimation consistency across frames.
Mehmet Hamza Erol, Arda Senocak, Jiu Feng +1 authors
AuM, a self-attention-free, state space model-based audio classifier, achieves performance comparable to or better than AST models, overcoming quadratic scaling issues.
Hao Wen, Zehuan Huang, Yaohui Wang +3 authors
A unified framework, Ouroboros3D, combines diffusion-based multi-view image generation and 3D reconstruction through a recursive diffusion process, improving geometric consistency and reducing data bias.
Eugene Choi, Arash Ahmadian, Matthieu Geist +2 authors
SRPO, a self-improving offline RLHF framework, achieves robustness to out-of-distribution tasks by optimizing a min-max objective that jointly enhances self-improvement and generative policies, leading to superior performance compared to DPO.
Elad Richardson, Yuval Alaluf, Ali Mahdavi-Amiri +1 authors
pOps framework uses Diffusion Prior model to train semantic operators in CLIP image embedding space for text-guided image generation.
Tao Yang, Yingmin Luo, Zhongang Qi +3 authors
A unified framework using a multi-modal large language model and data-driven methods for layout generation achieves state-of-the-art performance on benchmarks and introduces new datasets for real-world design tasks.
Xiefan Guo, Jinlin Liu, Miaomiao Cui +1 authors
I4VGen enhances text-to-video generation through image synthesis and guided video synthesis, using robust techniques to produce high-quality, realistic videos.
Tero Karras, Miika Aittala, Tuomas Kynkäänniemi +3 authors
Using smaller, less-trained versions of the model to guide generation leads to better prompt alignment and higher image quality without reducing variation in diffusion models.
Trung Dang, David Aponte, Dung Tran +1 authors
A novel autoregressive language model approach, LiveSpeech, enables low-latency text-to-speech by optimizing token prediction and parallel codebook processing.
Jiatao Gu, Ying Shen, Shuangfei Zhai +3 authors
Kaleido enhances diffusion model diversity by integrating autoregressive latent priors, generating varied and high-quality images from textual descriptions.
Edward Hughes, Michael Dennis, Jack Parker-Holder +5 authors
In recent years there has been a tremendous surge in the general capabilities of AI systems, mainly fuelled by training foundation models on internetscale data. Nevertheless, the creation of openended, ever self-improving AI remains elusive. In this position paper, we argue that the ingredients are now in place to achieve openendedness in AI systems with respect to a human observer. Furthermore, we claim that such open-endedness is an essential property of any artificial superhuman intelligence (ASI). We begin by providing a concrete formal definition of open-endedness through the lens of novelty and learnability. We then illustrate a path towards ASI via open-ended systems built on top of foundation models, capable of making novel, humanrelevant discoveries. We conclude by examining the safety implications of generally-capable openended AI. We expect that open-ended foundation models will prove to be an increasingly fertile and safety-critical area of research in the near future.
Jonathan Cook, Chris Lu, Edward Hughes +2 authors
Reinforcement learning agents can accumulate cultural knowledge and skills through social learning and generational training, outperforming single-lifetime training.
Haiyu Zhang, Xinyuan Chen, Yaohui Wang +3 authors
A novel 4D generation pipeline, 4Diffusion, uses a unified diffusion model with a learnable motion module to generate spatial-temporally consistent 4D content from monocular video, and introduces a 4D-aware Score Distillation Sampling loss and anchor loss to improve performance.
北京市昌平区探索星信息技术及软件开发工作室
京ICP备2026059466号