Reinforcement Pre-Training
Qingxiu Dong, Li Dong, Yao Tang +4 authors
Reinforcement Pre-Training (RPT) improves language model accuracy through reinforcement learning and offers a scalable method for leveraging text data for general-purpose RL.
Trends · 研究趋势
数据来自 Hugging Face 论文的 AI 提取关键词,按月统计研究方向的增长与热度。
Qingxiu Dong, Li Dong, Yao Tang +4 authors
Reinforcement Pre-Training (RPT) improves language model accuracy through reinforcement learning and offers a scalable method for leveraging text data for general-purpose RL.
NCP Team, Jiaqi Cao, Chiyu Chen +25 authors
NCP-ArchPreview is a large latent-space language model that jointly trains next-token and next-concept prediction to improve pretraining efficiency and downstream performance.
Meituan LongCat Team, Bin Xiao, Chao Wang +86 authors
Discrete Native Autoregressive framework enables unified multimodal processing by representing diverse modalities in a shared discrete space through a novel visual transformer architecture.
NextStep Team, Chunrui Han, Guopeng Li +47 authors
NextStep-1, a 14B autoregressive model with a 157M flow matching head, achieves state-of-the-art performance in text-to-image generation and image editing by processing discrete text tokens and continuous image tokens.
Subham Sekhar Sahoo, Lingjie Chen, Khiem Pham +14 authors
Diffusion-augmented autoregressive language models use parallel token sampling via distilled diffusion weights and a specialized sampler to accelerate inference without quality loss or draft models.
Yufeng Cui, Honghao Chen, Haoge Deng +20 authors
Emu3.5, a large-scale multimodal world model, predicts next states in vision and language, enhanced with reinforcement learning and Discrete Diffusion Adaptation for efficient inference, achieving strong performance in various multimodal tasks.
Shengbang Tong, David Fan, John Nguyen +18 authors
Controlled multimodal pretraining experiments reveal key insights about unified visual representations, data complementarity, world modeling emergence, and efficient scaling through mixture-of-experts architectures.
Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo +22 authors
Emu3, a transformer-based multimodal model trained exclusively with next-token prediction, outperforms existing diffusion and compositional models in generation and perception tasks.
Zhenghao Lin, Zhibin Gou, Yeyun Gong +8 authors
Rho-1, a novel language model using Selective Language Modeling, improves efficiency and performance by selectively training on useful tokens rather than all tokens in the corpus.
Peize Sun, Yi Jiang, Shoufa Chen +4 authors
LlamaGen applies the next-token prediction paradigm from large language models to image generation, achieving state-of-the-art performance with various model sizes and conditions.
Machel Reid, Nikolay Savinov, Denis Teplyashin +668 authors
Gemini 1.5 Pro, a multimodal mixture-of-experts model, achieves near-perfect recall and state-of-the-art performance in long-context tasks, including QA and ASR, surpassing previous models in token prediction and translation capabilities.
Dongchao Yang, Jinchuan Tian, Xu Tan +9 authors
UniAudio, a unified language model system, generates various audio types effectively using tokenization, sequence concatenation, and next-token prediction, with a multi-scale Transformer model handling extended sequences.
北京市昌平区探索星信息技术及软件开发工作室
京ICP备2026059466号