TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

606 篇论文 · 按点赞排序

543

Impossible Videos

Zechen Bai, Hai Ci, Mike Zheng Shou

IPV-Bench evaluates video generation and understanding models on creating and interpreting impossible videos, highlighting their limitations and guiding future advancements.

61video generation modelsprompt followingHF ↗arXiv ↗
544

Baichuan-Omni-1.5 Technical Report

Yadong Li, Jun Liu, Tao Zhang +90 authors

Baichuan-Omni-1.5 is an omni-modal model with end-to-end audio generation, featuring a comprehensive data pipeline, audio-tokenizer, and multi-stage training strategy for superior performance across multimodal tasks.

61omni-modal modelaudio-tokenizerHF ↗arXiv ↗
549

Antidistillation Sampling

Yash Savani, Asher Trockman, Zhili Feng +4 authors

Antidistillation sampling modifies a model's next-token probability distribution to disrupt the generation of reasoning traces for distillation without affecting model performance.

60antidistillation samplingnext-token probability distributionHF ↗arXiv ↗
555

Step-Audio-R1 Technical Report

Fei Tian, Xiangyu Tony Zhang, Yuxin Zhang +14 authors

Step-Audio-R1, using the Modality-Grounded Reasoning Distillation framework, achieves strong reasoning capabilities in audio, outperforming previous models and demonstrating the transferability of reasoning across modalities.

60reasoning modelschain-of-thought deliberationHF ↗arXiv ↗
558

Fino1: On the Transferability of Reasoning Enhanced LLMs to Finance

Lingfei Qian, Weipeng Zhou, Yan Wang +3 authors

A study evaluates 16 large language models on complex financial tasks, finding that domain-specific CoT fine-tuning and reinforcement learning improve performance and highlight the need for further research on long-context and multi-table reasoning.

59large language modelsfinancial reasoningHF ↗arXiv ↗
561

Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning

Chris, Yichen Wei, Yi Peng +10 authors

Skywork R1V2 enhances multimodal reasoning through a hybrid reinforcement learning approach that balances reward-model guidance and rule-based strategies, improving training efficiency with the Selective Sample Buffer mechanism and mitigating visual hallucinations.

59hybrid reinforcement learningreward-model guidanceHF ↗arXiv ↗
563

Gemma 3 Technical Report

Gemma Team, Aishwarya Kamath, Johan Ferret +213 authors

Gemma 3 introduces vision capabilities, broader language coverage, and extended context length, featuring an optimized architecture and post-training enhancements to outperform previous versions.

58multimodal modelsvision understandingHF ↗arXiv ↗
565

Chain-of-Retrieval Augmented Generation

Liang Wang, Haonan Chen, Nan Yang +3 authors

CoRAG, a multi-step retrieval and reasoning approach, enhances RAG models by dynamically refining queries and using rejection sampling to improve performance, especially in multi-hop question answering.

58RAG modelsCoRAGHF ↗arXiv ↗
568

Magma: A Foundation Model for Multimodal AI Agents

Jianwei Yang, Reuben Tan, Qianhui Wu +10 authors

Magma is a multimodal foundation model with both verbal intelligence and spatial-temporal intelligence, trained on diverse datasets to perform agentic tasks like UI navigation and robotic manipulation, outperforming specialized models.

58vision-language modelsspatial-temporal intelligenceHF ↗arXiv ↗
19 / 21

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号