TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

606 篇论文 · 按点赞排序

273

SemanticGen: Video Generation in Semantic Space

Jianhong Bai, Xiaoshi Wu, Xintao Wang +9 authors

SemanticGen addresses slow convergence and computational costs in video generation by using a two-stage diffusion model approach that first generates semantic features and then VAE latents, leading to faster convergence and high-quality results.

95VAE spaceVAE decoderHF ↗arXiv ↗
276

A Survey on Latent Reasoning

Rui-Jie Zhu, Tianhao Peng, Tianhao Cheng +30 authors

Latent reasoning in Large Language Models (LLMs) performs multi-step inference in continuous hidden states, enhancing reasoning capabilities without token-level supervision, and includes methodologies like activation-based recurrence and infinite-depth reasoning via masked diffusion models.

95chain-of-thought (CoT)latent reasoningHF ↗arXiv ↗
277

DeepSeek-OCR: Contexts Optical Compression

Haoran Wei, Yaofeng Sun, Yukun Li

DeepSeek-OCR uses optical 2D mapping to compress long contexts, achieving high OCR precision with reduced vision tokens and demonstrating practical value in document processing.

95DeepSeek-OCRDeepEncoderHF ↗arXiv ↗
284

Table-R1: Inference-Time Scaling for Table Reasoning

Zheyuan Yang, Lyuhao Chen, Arman Cohan +1 authors

Two post-training strategies, distillation and RLVR, enable inference-time scaling in table reasoning tasks, resulting in a model (Table-R1-Zero) that matches GPT-4.1's performance using fewer parameters and shows strong generalization.

93distillationreinforcement learningHF ↗arXiv ↗
288

Deep Think with Confidence

Yichao Fu, Xuewei Wang, Yuandong Tian +1 authors

DeepConf enhances reasoning efficiency and performance by filtering low-quality reasoning traces using model-internal confidence signals, achieving high accuracy and reducing token generation.

92Deep Think with ConfidenceDeepConfHF ↗arXiv ↗
289

Beyond Transcription: Mechanistic Interpretability in ASR

Neta Glazer, Yael Segal-Feldman, Hilit Segev +6 authors

Interpretability methods like logit lens, linear probing, and activation patching are applied to ASR to uncover internal dynamics, repetition hallucinations, and semantic biases, enhancing model transparency and robustness.

92logit lenslinear probingHF ↗arXiv ↗
290

GEM: A Gym for Agentic LLMs

Zichen Liu, Anya Sims, Keyu Duan +16 authors

GEM, an open-source environment simulator, facilitates experience-based learning for large language models by providing a standardized framework and diverse environments for training and benchmarking reinforcement learning algorithms.

92large language modelsexperience-based learningHF ↗arXiv ↗
292

Tree Search for LLM Agent Reinforcement Learning

Yuxiang Ji, Ziyu Ma, Yong Wang +3 authors

Tree-based Group Relative Policy Optimization (Tree-GRPO) enhances reinforcement learning for large language models by using tree search to improve rollouts and estimate grouped relative advantages, outperforming chain-based methods.

92reinforcement learninglarge language modelsHF ↗arXiv ↗
295

Yume: An Interactive World Generation Model

Xiaofeng Mao, Shaoheng Lin, Zhen Li +7 authors

A framework for generating and exploring interactive, high-fidelity video worlds from images using a Masked Video Diffusion Transformer, advanced sampling techniques, and model acceleration.

92camera motion quantizationMasked Video Diffusion TransformerHF ↗arXiv ↗
297

Tensor Product Attention Is All You Need

Yifan Zhang, Yifeng Liu, Huizhuo Yuan +4 authors

T6, a new Transformer architecture using Tensor Product Attention, improves language model performance and memory efficiency, allowing for longer sequence processing.

91Tensor Product AttentionTPAHF ↗arXiv ↗
300

Unveiling Intrinsic Dimension of Texts: from Academic Abstract to Creative Story

Vladislav Pedashenko, Laida Kushnareva, Yana Khassan Nibal +5 authors

The study explores intrinsic dimension in large language models through cross-encoder analysis, linguistic features, and sparse autoencoders, revealing its independence from entropy, genre-specific stratification, and causal features related to text type.

91intrinsic dimensioncross-encoder analysisHF ↗arXiv ↗
10 / 21

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号