TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Sep 8 – Sep 14, 2025

50 篇论文 · 按点赞排序

32

Reinforced Visual Perception with Tools

Zetong Zhou, Dongping Chen, Zixian Ma +6 authors

ReVPT enhances multi-modal LLMs' visual reasoning capabilities using reinforcement learning, achieving state-of-the-art performance on visual benchmarks.

32LLMsvision modelsHF ↗arXiv ↗
35

P3-SAM: Native 3D Part Segmentation

Changfeng Ma, Yang Li, Xinhao Yan +7 authors

P3-SAM, a native 3D point-promptable part segmentation model, achieves precise and robust segmentation of complex 3D objects using a feature extractor, multiple segmentation heads, and an IoU predictor.

263D point-promptable part segmentationP3-SAMHF ↗arXiv ↗
36

DivMerge: A divergence-based model merging method for multi-tasking

Touayouch Brahim, Fosse Loïc, Damnati Géraldine +1 authors

Multi-task learning (MTL) is often achieved by merging datasets before fine-tuning, but the growing availability of fine-tuned models has led to new approaches such as model merging via task arithmetic. A major challenge in this setting is task interference, which worsens as the number of tasks increases. We propose a method that merges models trained on different tasks into a single model, maintaining strong performance across all tasks. Our approach leverages Jensen-Shannon divergence to guide the merging process without requiring additional labelled data, and automatically balances task importance. Unlike existing methods, our approach remains robust as the number of tasks grows and consistently outperforms prior work.

26HF ↗arXiv ↗
37

Causal Attention with Lookahead Keys

Zhuoqing Song, Peng Sun, Huizhuo Yuan +1 authors

CASTLE, an attention mechanism that updates keys with future context while maintaining autoregressive properties, outperforms standard causal attention in language modeling.

22causal attentionQKVHF ↗arXiv ↗
39

Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning

Yuyao Ge, Shenghua Liu, Yiwei Wang +6 authors

Contrastive Attention Refinement for Visual Enhancement (CARVE) improves VLM performance by extracting task-relevant visual signals through attention contrasting, addressing issues with visual complexity and attention mechanisms.

20Vision-Language Modelsattention patternsHF ↗arXiv ↗
43

Interleaving Reasoning for Better Text-to-Image Generation

Wenxuan Huang, Shuang Chen, Zheyong Xie +15 authors

Interleaving Reasoning Generation (IRG) framework alternates between text-based thinking and image synthesis to improve Text-to-Image generation, achieving state-of-the-art performance and enhanced visual quality.

16Interleaving Reasoning GenerationIRGHF ↗arXiv ↗
46

Hunyuan-MT Technical Report

Mao Zheng, Zheng Li, Bingxin Qu +4 authors

Hunyuan-MT-7B and Hunyuan-MT-Chimera-7B are multilingual translation models that outperform existing models, especially in translating between Mandarin and minority languages, through a combination of pre-training, supervised fine-tuning, and reinforcement learning.

15multilingual translationbidirectional translationHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号