TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Apr 7 – Apr 13, 2025

50 篇论文 · 按点赞排序

32

OmniCaptioner: One Captioner to Rule Them All

Yiting Lu, Jiakang Yuan, Zhen Li +15 authors

OmniCaptioner generates detailed captions across various visual domains, enhancing visual reasoning with LLMs, improving image generation tasks, and enabling efficient supervised fine-tuning.

21visual captioning frameworklow-level pixel informationHF ↗arXiv ↗
36

Self-Steering Language Models

Gabriel Grand, Joshua B. Tenenbaum, Vikash K. Mansinghka +2 authors

DisCIPL, a method combining a Planner and Follower model, enables efficient and verifiable reasoning in language models by generating task-specific inference programs.

19test-time reasoninglanguage modelsHF ↗arXiv ↗
45

LiveVQA: Live Visual Knowledge Seeking

Mingyang Fu, Yuyang Peng, Benlin Liu +2 authors

Evaluation of various MLLMs on LiveVQA, a dataset of visual questions with up-to-date visual knowledge, shows that advanced visual reasoning is essential for complex multi-hop questions.

15MLLMsGPT-4oHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号