TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

606 篇论文 · 按点赞排序

302

LLaDA2.0: Scaling Up Diffusion Language Models to 100B

Tiwei Bie, Maosong Cao, Kun Chen +28 authors

LLaDA2.0 converts auto-regressive models into discrete diffusion large language models with a novel training scheme, achieving superior performance and efficiency at scale.

90discrete diffusion large language modelsdLLMHF ↗arXiv ↗
305

ZClip: Adaptive Spike Mitigation for LLM Pre-Training

Abhay Kumar, Louis Owen, Nilabhra Roy Chowdhury +1 authors

ZClip is an adaptive gradient clipping algorithm that uses z-score-based anomaly detection to dynamically adjust clipping thresholds and prevent large gradient spikes during LLM training.

90gradient instabilityloss spikesHF ↗arXiv ↗
306

Video-T1: Test-Time Scaling for Video Generation

Fangfu Liu, Hanyang Wang, Yimo Cai +3 authors

Test-Time Scaling (TTS) in video generation improves video quality by adaptively sampling from noise space with feedback mechanisms, particularly demonstrated with the Tree-of-Frames method.

90Test-Time Scaling (TTS)video generationHF ↗arXiv ↗
309

Flow-GRPO: Training Flow Matching Models via Online RL

Jie Liu, Gongye Liu, Jiajun Liang +6 authors

Flow-GRPO combines online reinforcement learning with flow matching models through an ODE-to-SDE conversion and denoising reduction, improving sampling efficiency and performance across text-to-image tasks.

89Flow-GRPOonline reinforcement learningHF ↗arXiv ↗
314

Step-DeepResearch Technical Report

Chen Hu, Haikuo Du, Heng Wang +64 authors

Step-DeepResearch, an end-to-end agent enhanced with a data synthesis strategy and progressive training, achieves expert-level capabilities in deep research scenarios, outperforming established models.

89Deep ResearchBrowseCompHF ↗arXiv ↗
317

Learning to Reason under Off-Policy Guidance

Jianhao Yan, Yafu Li, Zican Hu +5 authors

LUFFY enhances zero-RL models with off-policy guidance, improving reasoning and generalization through balanced imitation and exploration.

88large reasoning modelsreinforcement learningHF ↗arXiv ↗
320

Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought

Yi Peng, Chris, Xiaokun Wang +12 authors

Skywork R1V extends large language models to multimodal reasoning with efficient transfer, enhanced visual-text alignment, and dynamic reasoning chain optimization, achieving competitive performance in various benchmarks.

87multimodal reasoning modelR1-series Large language modelsHF ↗arXiv ↗
321

BitNet b1.58 2B4T Technical Report

Shuming Ma, Hongyu Wang, Shaohan Huang +5 authors

BitNet b1.58 2B4T, a 1-bit Large Language Model with 2 billion parameters, matches the performance of full-precision models while improving computational efficiency.

87BitNetLarge Language ModelHF ↗arXiv ↗
322

Can LLMs Predict Their Own Failures? Self-Awareness via Internal Circuits

Amirhosein Ghasemabadi, Di Niu

Large language models (LLMs) generate fluent and complex outputs but often fail to recognize their own mistakes and hallucinations. Existing approaches typically rely on external judges, multi-sample consistency, or text-based self-critique, which incur additional compute or correlate weakly with true correctness. We ask: can LLMs predict their own failures by inspecting internal states during inference? We introduce Gnosis, a lightweight self-awareness mechanism that enables frozen LLMs to perform intrinsic self-verification by decoding signals from hidden states and attention patterns. Gnosis passively observes internal traces, compresses them into fixed-budget descriptors, and predicts correctness with negligible inference cost, adding only ~5M parameters and operating independently of sequence length. Across math reasoning, open-domain question answering, and academic knowledge benchmarks, and over frozen backbones ranging from 1.7B to 20B parameters, Gnosis consistently outperforms strong internal baselines and large external judges in both accuracy and calibration. Moreover, it generalizes zero-shot to partial generations, enabling early detection of failing trajectories and compute-aware control. These results show that reliable correctness cues are intrinsic to generation process and can be extracted efficiently without external supervision.

87HF ↗arXiv ↗
324

KORMo: Korean Open Reasoning Model for Everyone

Minjun Kim, Hyeonseok Lim, Hangyeol Yoo +10 authors

A large-scale investigation into constructing a fully open bilingual LLM for Korean using synthetic data demonstrates that such data can sustain pretraining and achieve performance comparable to multilingual baselines.

87large language modelLLMHF ↗arXiv ↗
330

Visual-RFT: Visual Reinforcement Fine-Tuning

Ziyu Liu, Zeyi Sun, Yuhang Zang +5 authors

Visual Reinforcement Fine-Tuning (Visual-RFT) enhances Large Vision-Language Models (LVLMs) through reinforcement learning with visual perception verifiable rewards, achieving competitive performance in various visual and reasoning tasks with limited data.

86Reinforcement Fine-Tuning (RFT)Visual Reinforcement Fine-Tuning (Visual-RFT)HF ↗arXiv ↗
11 / 21

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号