TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Feb 3 – Feb 9, 2025
本周最热260

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model

Loubna Ben Allal, Anton Lozhkov, Elie Bakouch +19 authors

SmolLM2, a small language model with 1.7 billion parameters, achieves strong performance through overtraining on diverse datasets, outperforming other recent small models.

large language modelssmall language modelsovertrainingdataset mixingHF ↗arXiv ↗

50 篇论文 · 按点赞排序

03

s1: Simple test-time scaling

Niklas Muennighoff, Zitong Yang, Weijia Shi +7 authors

The use of budget forcing during test time improves reasoning performance in language models, as demonstrated by the s1 model which outperforms the o1-preview on competition math questions.

128test-time scalinglanguage modelingHF ↗arXiv ↗
04

The Differences Between Direct Alignment Algorithms are a Blur

Alexey Gorbatovski, Boris Shaposhnikov, Viacheslav Sinii +2 authors

Direct Alignment Algorithms improve language model alignment by introducing a supervised fine-tuning phase and adjusting preference optimization strength, showing that ranking objectives are crucial for performance.

113Direct Alignment AlgorithmsRLHFHF ↗arXiv ↗
06

LIMO: Less is More for Reasoning

Yixin Ye, Zhen Huang, Yang Xiao +3 authors

LIMO, a new model, achieves high mathematical reasoning performance using minimal training data, challenging the notion that extensive datasets are necessary for complex reasoning.

63LIMOLIMO HypothesisHF ↗arXiv ↗
07

Process Reinforcement through Implicit Rewards

Ganqu Cui, Lifan Yuan, Zefan Wang +20 authors

PRIME leverages implicit process rewards to improve the reinforcement learning of large language models, achieving better performance with less data compared to traditional methods.

62dense process rewardssparse outcome-level rewardsHF ↗arXiv ↗
09

Demystifying Long Chain-of-Thought Reasoning in LLMs

Edward Yeo, Yuxuan Tong, Morry Niu +2 authors

Investigation into long chains-of-thought reasoning in large language models reveals the critical role of training compute, reward shaping, and verifiable reward signals in enabling and measuring this capability.

57large language modelslong chains-of-thoughtHF ↗arXiv ↗
14

Reward-Guided Speculative Decoding for Efficient LLM Reasoning

Baohao Liao, Yuhui Xu, Hanze Dong +5 authors

Reward-Guided Speculative Decoding improves inference efficiency in large language models by combining a draft and target model with a reward-based strategy, achieving better performance and cost savings.

39Reward-Guided Speculative Decodingspeculative decodingHF ↗arXiv ↗
19

DynVFX: Augmenting Real Videos with Dynamic Content

Danah Yatim, Rafail Fridman, Omer Bar-Tal +1 authors

A zero-shot, training-free method using text-to-video diffusion transformer and Vision Language Model synthesizes and integrates dynamic content into real-world videos based on text instructions.

29text-to-video diffusion transformerVision Language ModelHF ↗arXiv ↗
22

Inverse Bridge Matching Distillation

Nikita Gushchin, David Li, Daniil Selikhanovych +3 authors

A novel distillation method accelerates diffusion bridge models for image-to-image translation tasks, improving inference speed and quality.

28diffusion bridge modelsinverse bridge matchingHF ↗arXiv ↗
27

Weak-to-Strong Diffusion with Reflection

Lichen Bai, Masashi Sugiyama, Zeke Xie

The Weak-to-Strong Diffusion (W2SD) framework enhances the alignment of learned distributions with real data distributions by leveraging differences between weak and strong models, leading to significant improvements in human preference, aesthetic quality, and prompt adherence across multiple modalities and architectures.

24diffusion generative modelsgradient score matchingHF ↗arXiv ↗
29

Scaling Embedding Layers in Language Models

Da Yu, Edith Cohen, Badih Ghazi +5 authors

SCONE extends language model performance by using contextualized n-gram embeddings, precomputed and stored off-accelerator, to minimize inference costs while maintaining fixed FLOPS.

24n-gram embeddingscontextualized representationHF ↗arXiv ↗
30

Scalable-Softmax Is Superior for Attention

Ken M. Nakanishi

Scalable-Softmax (SSMax) enhances Transformer-based models by improving attention distribution and long-context performance, resolving issues posed by the Softmax function.

24Softmaxattention scoresHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号