TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

本月最热234

Feature-Level Insights into Artificial Text Detection with Sparse Autoencoders

Kristian Kuznetsov, Laida Kushnareva, Polina Druzhinina +5 authors

Sparse Autoencoders enhance interpretability in Artificial Text Detection by extracting distinctive features from LLM outputs, providing insights into differences from human-written texts.

Sparse Autoencodersresidual streaminterpretable featuresdomain-specific statisticsHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

Qwen2.5-Omni Technical Report

Jin Xu, Zhifang Guo, Jinzheng He +11 authors

Qwen2.5-Omni is a multimodal model that processes text, images, audio, and video in a streaming fashion and generates text and speech using a dual-track architecture, achieving state-of-the-art performance on multimodal benchmarks.

174block-wise processingTMRoPE (Time-aligned Multimodal RoPE)HF ↗arXiv ↗
04

Transformers without Normalization

Jiachen Zhu, Xinlei Chen, Kaiming He +2 authors

Dynamic Tanh (DyT) replaces normalization layers in Transformers, achieving equivalent or superior performance without hyperparameter tuning across various tasks.

171Normalization layersDynamic TanhHF ↗arXiv ↗
05

RWKV-7 "Goose" with Expressive Dynamic State Evolution

Bo Peng, Ruichong Zhang, Daniel Goldstein +12 authors

RWKV-7 "Goose" achieves state-of-the-art performance in multilingual tasks with optimal memory and inference efficiency, exceeding Transformer capabilities in complexity.

155sequence modeling architecturedelta ruleHF ↗arXiv ↗
06

DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Qiying Yu, Zheng Zhang, Ruofei Zhu +32 authors

The DAPO algorithm, which includes decoupled clip and dynamic sampling policy optimization, enables open-source, high-performance reinforcement learning training for large-scale language models, enhancing reproducibility in the field.

148LLMsreinforcement learningHF ↗arXiv ↗
07

ReCamMaster: Camera-Controlled Generative Rendering from A Single Video

Jianhong Bai, Menghan Xia, Xiao Fu +8 authors

ReCamMaster uses pre-trained text-to-video models to render dynamic scenes of input videos from novel camera trajectories, leveraging a unique video conditioning mechanism and a custom dataset created with Unreal Engine 5.

148camera-controlled generative video re-renderingpre-trained text-to-video modelsHF ↗arXiv ↗
09

RuCCoD: Towards Automated ICD Coding in Russian

Aleksandr Nesterov, Andrey Sakhovskiy, Ivan Sviridov +5 authors

Experiments on a new Russian-language ICD coding dataset using models like BERT, LLaMA with LoRA, and RAG show significant accuracy improvements in automated clinical coding compared to manual annotations.

133BERTLLaMAHF ↗arXiv ↗
12

START: Self-taught Reasoner with Tools

Chengpeng Li, Mingfeng Xue, Zhenru Zhang +7 authors

START integrates external tools into large reasoning models to enhance capabilities, using techniques like Hint-infer and Hint Rejection Sampling Fine-Tuning, achieving high performance across various benchmarks.

113Large reasoning modelsChain-of-thoughtHF ↗arXiv ↗
14

Survey on Evaluation of LLM-based Agents

Asaf Yehudai, Lilach Eden, Alan Li +5 authors

This survey analyzes evaluation methodologies for large language model-based agents, covering fundamental capabilities, application-specific benchmarks, and generalist agents, highlighting trends and gaps in the field.

97LLM-based agentsautonomous systemsHF ↗arXiv ↗
19

Video-T1: Test-Time Scaling for Video Generation

Fangfu Liu, Hanyang Wang, Yimo Cai +3 authors

Test-Time Scaling (TTS) in video generation improves video quality by adaptively sampling from noise space with feedback mechanisms, particularly demonstrated with the Tree-of-Frames method.

90Test-Time Scaling (TTS)video generationHF ↗arXiv ↗
22

Visual-RFT: Visual Reinforcement Fine-Tuning

Ziyu Liu, Zeyi Sun, Yuhang Zang +5 authors

Visual Reinforcement Fine-Tuning (Visual-RFT) enhances Large Vision-Language Models (LVLMs) through reinforcement learning with visual perception verifiable rewards, achieving competitive performance in various visual and reasoning tasks with limited data.

86Reinforcement Fine-Tuning (RFT)Visual Reinforcement Fine-Tuning (Visual-RFT)HF ↗arXiv ↗
24

EuroBERT: Scaling Multilingual Encoders for European Languages

Nicolas Boizard, Hippolyte Gisserot-Boukhlef, Duarte M. Alves +16 authors

EuroBERT, a family of multilingual encoders covering European and global languages, outperforms existing models across various tasks and supports long sequences, surpassing traditional bidirectional encoders.

81bidirectional encoder modelsgenerative decoder-only modelsHF ↗arXiv ↗
25

Video-R1: Reinforcing Video Reasoning in MLLMs

Kaituo Feng, Kaixiong Gong, Bohao Li +5 authors

Video-R1, leveraging rule-based reinforcement learning and temporal information, enhances video reasoning in multimodal large language models using a combination of video and image data.

79rule-based reinforcement learningRLHF ↗arXiv ↗
30

YuE: Scaling Open Foundation Models for Long-Form Music Generation

Ruibin Yuan, Hanfeng Lin, Shuyue Guo +54 authors

YuE, a family of open foundation models based on LLaMA2, can generate long-form music with aligned lyrics, coherent structure, and appropriate accompaniment using innovative techniques in next-token prediction, conditioning, and pre-training.

75track-decoupled next-token predictionstructural progressive conditioningHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号