TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Feb 10 – Feb 16, 2025
本周最热193

The Stochastic Parrot on LLM's Shoulder: A Summative Assessment of Physical Concept Understanding

Mo Yu, Lemao Liu, Junjie Wu +5 authors

A study investigates whether large language models understand physical concepts through a grid-based task, showing that they significantly underperform compared to humans and highlighting the stochastic parrot phenomenon.

LLMsStochastic ParrotPhysiCogrid-format inputsHF ↗arXiv ↗

50 篇论文 · 按点赞排序

05

Expect the Unexpected: FailSafe Long Context QA for Finance

Kiran Kamble, Melisa Russak, Dmytro Mozolevskyi +3 authors

FailSafeQA evaluates the robustness and context-awareness of large language models in financial applications through domain expertise, query completeness, linguistic accuracy, and document relevance challenges.

132LLMlong-context financial benchmarkHF ↗arXiv ↗
06

Goku: Flow Based Video Generative Foundation Models

Shoufa Chen, Chongjian Ge, Yuqi Zhang +19 authors

Goku, a state-of-the-art family of joint image-and-video generation models using rectified flow Transformers, sets new benchmarks in text-to-image and text-to-video tasks.

106rectified flow Transformersimage-and-video generation modelsHF ↗arXiv ↗
09

Competitive Programming with Large Reasoning Models

OpenAI, Ahmed El-Kishky, Alexander Wei +22 authors

General-purpose reinforcement learning applied to large language models outperforms domain-specific systems in complex coding and reasoning tasks, achieving top results in competitions without hand-crafted strategies.

69reinforcement learninglarge language modelsHF ↗arXiv ↗
11

Fino1: On the Transferability of Reasoning Enhanced LLMs to Finance

Lingfei Qian, Weipeng Zhou, Yan Wang +3 authors

A study evaluates 16 large language models on complex financial tasks, finding that domain-specific CoT fine-tuning and reinforcement learning improve performance and highlight the need for further research on long-context and multi-table reasoning.

59large language modelsfinancial reasoningHF ↗arXiv ↗
16

Distillation Scaling Laws

Dan Busbridge, Amitis Shidani, Floris Weers +3 authors

The study presents a scaling law to optimize compute allocation for model distillation, showing conditions under which distillation outperforms supervised pretraining.

48distillation scaling lawdistilled modelHF ↗arXiv ↗
18

QuEST: Stable Training of LLMs with 1-Bit Weights and Activations

Andrei Panferov, Jiale Chen, Soroush Tabesh +3 authors

QuEST, a new method for quantization-aware training, achieves better accuracy at lower model size with weights and activations in 4-bits or less, even stable training with 1-bit, by improving quantization accuracy and the gradient estimator.

44Quantization-Aware Training (QAT)QuESTHF ↗arXiv ↗
24

The Curse of Depth in Large Language Models

Wenfang Sun, Xinyuan Song, Pengxiang Li +3 authors

LayerNorm Scaling addresses the Curse of Depth in LLMs by scaling variance inversely with layer depth, improving pre-training and fine-tuning effectiveness of deep layers.

40Curse of DepthLarge Language ModelsHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号