TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

本月最热312

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Jinguo Zhu, Weiyun Wang, Zhe Chen +44 authors

InternVL3 is a multimodal pre-trained language model that jointly learns from both multimodal data and text, improving performance and scalability through advanced techniques and setting a new state-of-the-art in multimodal tasks.

multimodal pre-traininglarge language modelmultimodal large language modelvariable visual position encodingHF ↗arXiv ↗

50 篇论文 · 按点赞排序

05

Towards Understanding Camera Motions in Any Video

Zhiqiu Lin, Siyuan Cen, Daniel Jiang +12 authors

CameraBench assesses and enhances camera motion understanding by providing a large-scale dataset, taxonomy, and benchmark for evaluating SfM and VLMs, emphasizing the need for both semantic and geometric information.

157Structure-from-Motion (SfM)Video-Language Models (VLMs)HF ↗arXiv ↗
06

Kimi-VL Technical Report

Kimi Team, Angang Du, Bohong Yin +89 authors

Kimi-VL, an efficient Mixture-of-Experts vision-language model, excels in multimodal reasoning, long-context understanding, and diverse vision-language tasks, achieving competitive performance with reduced computational cost.

143Mixture-of-Experts (MoE)vision-language model (VLM)HF ↗arXiv ↗
09

MoCha: Towards Movie-Grade Talking Character Synthesis

Cong Wei, Bo Sun, Haoyu Ma +10 authors

MoCha generates realistic talking character animations from speech and text using a speech-video attention mechanism and joint training on speech-labeled and text-labeled data, enabling multi-character conversations and superior realism.

141MoChaspeech-video window attention mechanismHF ↗arXiv ↗
12

TTRL: Test-Time Reinforcement Learning

Yuxin Zuo, Kaiyan Zhang, Shang Qu +7 authors

Test-Time Reinforcement Learning (TTRL) enhances Large Language Models (LLMs) using unlabeled data through reinforcement learning, improving performance across tasks.

123Reinforcement Learning (RL)Large Language Models (LLMs)HF ↗arXiv ↗
15

One-Minute Video Generation with Test-Time Training

Karan Dalal, Daniel Koceja, Gashon Hussein +12 authors

Test-Time Training (TTT) layers enable pre-trained Transformers to generate coherent one-minute videos from text storyboards, outperforming alternatives like Mamba~2 and Gated DeltaNet.

110self-attention layersMamba layersHF ↗arXiv ↗
17

CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training

Shizhe Diao, Yu Yang, Yonggan Fu +12 authors

Pre-training datasets are typically collected from web content and lack inherent domain divisions. For instance, widely used datasets like Common Crawl do not include explicit domain labels, while manually curating labeled datasets such as The Pile is labor-intensive. Consequently, identifying an optimal pre-training data mixture remains a challenging problem, despite its significant benefits for pre-training performance. To address these challenges, we propose CLustering-based Iterative Data Mixture Bootstrapping (CLIMB), an automated framework that discovers, evaluates, and refines data mixtures in a pre-training setting. Specifically, CLIMB embeds and clusters large-scale datasets in a semantic space and then iteratively searches for optimal mixtures using a smaller proxy model and a predictor. When continuously trained on 400B tokens with this mixture, our 1B model exceeds the state-of-the-art Llama-3.2-1B by 2.0%. Moreover, we observe that optimizing for a specific domain (e.g., Social Sciences) yields a 5% improvement over random sampling. Finally, we introduce ClimbLab, a filtered 1.2-trillion-token corpus with 20 clusters as a research playground, and ClimbMix, a compact yet powerful 400-billion-token dataset designed for efficient pre-training that delivers superior performance under an equal token budget. We analyze the final data mixture, elucidating the characteristics of an optimal data mixture. Our data is available at: https://research.nvidia.com/labs/lpr/climb/

98CLIMBsemantic spaceHF ↗arXiv ↗
21

ZClip: Adaptive Spike Mitigation for LLM Pre-Training

Abhay Kumar, Louis Owen, Nilabhra Roy Chowdhury +1 authors

ZClip is an adaptive gradient clipping algorithm that uses z-score-based anomaly detection to dynamically adjust clipping thresholds and prevent large gradient spikes during LLM training.

90gradient instabilityloss spikesHF ↗arXiv ↗
22

Learning to Reason under Off-Policy Guidance

Jianhao Yan, Yafu Li, Zican Hu +5 authors

LUFFY enhances zero-RL models with off-policy guidance, improving reasoning and generalization through balanced imitation and exploration.

88large reasoning modelsreinforcement learningHF ↗arXiv ↗
23

BitNet b1.58 2B4T Technical Report

Shuming Ma, Hongyu Wang, Shaohan Huang +5 authors

BitNet b1.58 2B4T, a 1-bit Large Language Model with 2 billion parameters, matches the performance of full-precision models while improving computational efficiency.

87BitNetLarge Language ModelHF ↗arXiv ↗
24

Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought

Yi Peng, Chris, Xiaokun Wang +12 authors

Skywork R1V extends large language models to multimodal reasoning with efficient transfer, enhanced visual-text alignment, and dynamic reasoning chain optimization, achieving competitive performance in various benchmarks.

87multimodal reasoning modelR1-series Large language modelsHF ↗arXiv ↗
29

DDT: Decoupled Diffusion Transformer

Shuai Wang, Zhi Tian, Weilin Huang +1 authors

A decoupled diffusion transformer improves performance and training speed in image generation by separating semantic extraction and high-frequency decoding.

77diffusion transformersdenoising stepsHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号