TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Apr 7 – Apr 13, 2025
本周最热212

SmolVLM: Redefining small and efficient multimodal models

Andrés Marafioti, Orr Zohar, Miquel Farré +14 authors

SmolVLM, a series of compact multimodal models, achieves high performance with minimal GPU memory usage, making efficient deployment on mobile and edge devices possible.

Large Vision-Language ModelsVLMsSmolVLMmultimodal modelsHF ↗arXiv ↗

50 篇论文 · 按点赞排序

03

Kimi-VL Technical Report

Kimi Team, Angang Du, Bohong Yin +89 authors

Kimi-VL, an efficient Mixture-of-Experts vision-language model, excels in multimodal reasoning, long-context understanding, and diverse vision-language tasks, achieving competitive performance with reduced computational cost.

143Mixture-of-Experts (MoE)vision-language model (VLM)HF ↗arXiv ↗
05

One-Minute Video Generation with Test-Time Training

Karan Dalal, Daniel Koceja, Gashon Hussein +12 authors

Test-Time Training (TTT) layers enable pre-trained Transformers to generate coherent one-minute videos from text storyboards, outperforming alternatives like Mamba~2 and Gated DeltaNet.

110self-attention layersMamba layersHF ↗arXiv ↗
06

Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought

Yi Peng, Chris, Xiaokun Wang +12 authors

Skywork R1V extends large language models to multimodal reasoning with efficient transfer, enhanced visual-text alignment, and dynamic reasoning chain optimization, achieving competitive performance in various benchmarks.

87multimodal reasoning modelR1-series Large language modelsHF ↗arXiv ↗
09

DDT: Decoupled Diffusion Transformer

Shuai Wang, Zhi Tian, Weilin Huang +1 authors

A decoupled diffusion transformer improves performance and training speed in image generation by separating semantic extraction and high-frequency decoding.

77diffusion transformersdenoising stepsHF ↗arXiv ↗
10

An Empirical Study of GPT-4o Image Generation Capabilities

Sixiang Chen, Jinbin Bai, Zhuoran Zhao +16 authors

An empirical study of GPT-4o's image generation capabilities across multiple tasks reveals its strengths and limitations compared to other models, highlighting the importance of architectural design and data scaling in unified generative frameworks.

64GANdiffusion modelsHF ↗arXiv ↗
13

Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving

Daoguang Zan, Zhirong Huang, Wei Liu +16 authors

A multilingual benchmark, Multi-SWE-bench, is introduced to evaluate Large Language Models across various programming languages and is complemented by a RL dataset, Multi-SWE-RL, to advance issue-resolving tasks.

49Large Language Models (LLMs)multilingual issue-resolving benchmarkHF ↗arXiv ↗
18

Rethinking Reflection in Pre-Training

Essential AI, Darsh J Shah, Peter Rushton +25 authors

Models exhibit self-correcting ability during pre-training by recognizing and addressing errors in their reasoning, a skill that improves over time.

37language modelself-reflectionHF ↗arXiv ↗
19

MM-IFEngine: Towards Multimodal Instruction Following

Shengyuan Ding, Shenxi Wu, Xiangyu Zhao +7 authors

MM-IFEngine generates high-quality image-instruction pairs for training Multi-modal Large Language Models, leading to improved performance in instruction-following tasks.

35Multi-modal Large Language ModelsMM-IFEngineHF ↗arXiv ↗
22

URECA: Unique Region Caption Anything

Sangbeom Lim, Junwan Kim, Heeji Yoon +2 authors

The URECA dataset and model address multi-granularity region captioning by ensuring unique and contextually grounded captions through a stage-wise pipeline and dynamic mask modeling.

35region-level captioningURECA datasetHF ↗arXiv ↗
23

MegaMath: Pushing the Limits of Open Math Corpora

Fan Zhou, Zengzhi Wang, Nikhil Ranjan +5 authors

MegaMath is an open dataset designed for math-centric LLM pre-training, combining high-quality web data, math-related code, and synthetic content to enhance diversity and quality.

35LLMsMegaMathHF ↗arXiv ↗
27

HoloPart: Generative 3D Part Amodal Segmentation

Yunhan Yang, Yuan-Chen Guo, Yukun Huang +5 authors

HoloPart, a diffusion-based model with local and global attention, addresses 3D part amodal segmentation by completing occluded parts and improving global shape consistency.

28diffusion-based modellocal attentionHF ↗arXiv ↗
29

Agentic Knowledgeable Self-awareness

Shuofei Qiao, Zhisong Qiu, Baochang Ren +8 authors

KnowSelf, a data-centric approach, enables LLM-based agents to autonomously regulate knowledge utilization and achieve optimal planning with minimal costs by switching between situations.

27LLMsagentic planningHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号