TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

606 篇论文 · 按点赞排序

333

Seedream 4.0: Toward Next-generation Multimodal Image Generation

Team Seedream, Yunpeng Chen, Yu Gao +47 authors

Seedream 4.0 is a high-performance multimodal image generation system that integrates text-to-image synthesis, image editing, and multi-image composition using a diffusion transformer and VAE, achieving state-of-the-art results with efficient training and inference.

85diffusion transformerVAEHF ↗arXiv ↗
336

Self-Rewarding Vision-Language Model via Reasoning Decomposition

Zongxia Li, Wenhao Yu, Chengsong Huang +8 authors

Vision-SR1 uses reinforcement learning to enhance visual reasoning in vision-language models by decomposing the process into visual perception and language reasoning stages, improving accuracy and reducing hallucinations.

85vision-language modelsvisual hallucinationsHF ↗arXiv ↗
337

The Invisible Leash: Why RLVR May Not Escape Its Origin

Fang Wu, Weihao Xuan, Ximing Lu +2 authors

Reinforcement Learning with Verifiable Rewards (RLVR) enhances precision but may limit exploration and discovery of new solutions, suggesting potential limits to its effectiveness in expanding reasoning capabilities.

85Reinforcement Learning with Verifiable RewardsRLVRHF ↗arXiv ↗
341

Soundwave: Less is More for Speech-Text Alignment in LLMs

Yuhao Zhang, Zhiheng Liu, Fan Bu +3 authors

Soundwave addresses the representation space gap and sequence length inconsistency in end-to-end speech large language models using an efficient training strategy and novel architecture, outperforming Qwen2-Audio with significantly less data.

85large language modelslarge-scale annotated dataHF ↗arXiv ↗
347

Parallel Scaling Law for Language Models

Mouxiang Chen, Binyuan Hui, Zeyu Cui +5 authors

Parallel scaling (ParScale) improves inference efficiency by reusing existing parameters and executing multiple transformations in parallel, offering superior performance with reduced memory and latency compared to parameter scaling.

83parallel computationparallel scalingHF ↗arXiv ↗
350

RAG-Anything: All-in-One RAG Framework

Zirui Guo, Xubin Ren, Lingrui Xu +2 authors

RAG-Anything is a unified framework that enhances multimodal knowledge retrieval by integrating cross-modal relationships and semantic matching, outperforming existing methods on complex benchmarks.

83Retrieval-Augmented GenerationRAGHF ↗arXiv ↗
351

Scaling Agent Learning via Experience Synthesis

Zhaorun Chen, Zhuokai Zhao, Kai Zhang +15 authors

DreamGym is a unified framework that synthesizes diverse experiences for scalable online RL training, improving agent performance and reducing real-world interactions.

83reinforcement learninglarge language modelHF ↗arXiv ↗
355

ExGRPO: Learning to Reason from Experience

Runzhe Zhan, Yafu Li, Zhi Wang +5 authors

ExGRPO, a framework that prioritizes valuable reasoning experiences, improves and stabilizes reinforcement learning from verifiable rewards for large language models.

83reinforcement learning from verifiable rewardsRLVRHF ↗arXiv ↗
360

MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Yilun Zhao, Lujing Xie, Haowei Zhang +16 authors

A comprehensive benchmark for evaluating foundation models in video understanding includes domain-specific questions and expert evaluations, highlighting the gap between current models and human expertise.

82multimodal foundation modelsexpert-level reasoningHF ↗arXiv ↗
12 / 21

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号