TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Mar 31 – Apr 6, 2025

50 篇论文 · 按点赞排序

31

Exploring the Effect of Reinforcement Learning on Video Understanding: Insights from SEED-Bench-R1

Yi Chen, Yuying Ge, Rui Wang +4 authors

A benchmark (SEED-Bench-R1) evaluates reinforcement learning versus supervised fine-tuning for multimodal large language models in video understanding, highlighting RL's data efficiency and superior performance but identifying limitations in logical coherence and visual cue processing.

38Chain of Thought (COT)Large Language Models (LLMs)HF ↗arXiv ↗
32

WikiVideo: Article Generation from Multiple Videos

Alexander Martin, Reno Kriz, William Gantt Walden +5 authors

WikiVideo proposes Collaborative Article Generation (CAG) to enhance high-level event summarization from videos by integrating r1-style reasoning and VideoLLM for better inferences compared to state-of-the-art VideoLLMs.

37retrieval-augmented generationRAGHF ↗arXiv ↗
36

Scaling Language-Free Visual Representation Learning

David Fan, Shengbang Tong, Jiachen Zhu +8 authors

Visual self-supervised learning matches language-supervised visual pretraining performance on VQA and vision benchmarks when both are trained on the same dataset and scaled appropriately.

33visual self-supervised learningcontrastive language-image pretrainingHF ↗arXiv ↗
38

Command A: An Enterprise-Ready Large Language Model

Team Cohere, Aakanksha, Arash Ahmadian +223 authors

Command A, a multilingual large language model, uses decentralized training with self-refinement and model merging to achieve efficient and top-performing Retrieval Augmented Generation for enterprise use.

31agent-optimisedmultilingual-capableHF ↗arXiv ↗
42

Z1: Efficient Test-time Scaling with Code

Zhaojian Yu, Yinghao Wu, Yilun Zhao +2 authors

A novel Shifted Thinking Window method trains LLMs on code-related reasoning trajectories to reduce excessive thinking tokens while maintaining performance and efficient test-time scaling.

27LLMstest-time computing scalingHF ↗arXiv ↗
47

Your ViT is Secretly an Image Segmentation Model

Tommie Kerssies, Niccolò Cavagnero, Alexander Hermans +5 authors

The Encoder-only Mask Transformer (EoMT) achieves state-of-the-art image segmentation accuracy by learning task-specific components through large-scale pre-training, while maintaining superior prediction speed compared to models with additional architectural complexity.

25Vision TransformersViTsHF ↗arXiv ↗
48

ReFeed: Multi-dimensional Summarization Refinement with Reflective Reasoning on Feedback

Taewon Yun, Jihwan Oh, Hyangsuk Min +4 authors

Summarization refinement faces challenges when extending to multi-dimension. In this paper, we introduce ReFeed, a powerful summarization refinement pipeline that enhances multiple dimensions through reflective reasoning on feedback. To achieve this, we release SumFeed-CoT, a large-scale Long-CoT-based dataset optimized for training a lightweight model with reflective reasoning. Our experiments reveal how the number of dimensions, feedback exposure, and reasoning policy influence refinement performance, highlighting reflective reasoning and simultaneously addressing multiple feedback is crucial to mitigate trade-off between dimensions. Furthermore, ReFeed is robust to noisy feedback and feedback order. Lastly, our finding emphasizes that creating data with a proper goal and guideline constitutes a fundamental pillar of effective reasoning. The dataset and model will be released.

25Long-CoT-based datasetreflective reasoningHF ↗arXiv ↗
50

Expanding RL with Verifiable Rewards Across Diverse Domains

Yi Su, Dian Yu, Linfeng Song +5 authors

Reinforcement learning with verifiable rewards extended to diverse domains using model-based soft scoring outperforms existing LLMs in free-form answer settings with reliable reward signals.

24reinforcement learning with verifiable rewardslarge language modelsHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号