TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Mar 10 – Mar 16, 2025
本周最热234

Feature-Level Insights into Artificial Text Detection with Sparse Autoencoders

Kristian Kuznetsov, Laida Kushnareva, Polina Druzhinina +5 authors

Sparse Autoencoders enhance interpretability in Artificial Text Detection by extracting distinctive features from LLM outputs, providing insights into differences from human-written texts.

Sparse Autoencodersresidual streaminterpretable featuresdomain-specific statisticsHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

Transformers without Normalization

Jiachen Zhu, Xinlei Chen, Kaiming He +2 authors

Dynamic Tanh (DyT) replaces normalization layers in Transformers, achieving equivalent or superior performance without hyperparameter tuning across various tasks.

171Normalization layersDynamic TanhHF ↗arXiv ↗
03

RuCCoD: Towards Automated ICD Coding in Russian

Aleksandr Nesterov, Andrey Sakhovskiy, Ivan Sviridov +5 authors

Experiments on a new Russian-language ICD coding dataset using models like BERT, LLaMA with LoRA, and RAG show significant accuracy improvements in automated clinical coding compared to manual annotations.

133BERTLLaMAHF ↗arXiv ↗
08

EuroBERT: Scaling Multilingual Encoders for European Languages

Nicolas Boizard, Hippolyte Gisserot-Boukhlef, Duarte M. Alves +16 authors

EuroBERT, a family of multilingual encoders covering European and global languages, outperforms existing models across various tasks and supports long sequences, surpassing traditional bidirectional encoders.

81bidirectional encoder modelsgenerative decoder-only modelsHF ↗arXiv ↗
11

YuE: Scaling Open Foundation Models for Long-Form Music Generation

Ruibin Yuan, Hanfeng Lin, Shuyue Guo +54 authors

YuE, a family of open foundation models based on LLaMA2, can generate long-form music with aligned lyrics, coherent structure, and appropriate accompaniment using innovative techniques in next-token prediction, conditioning, and pre-training.

77track-decoupled next-token predictionstructural progressive conditioningHF ↗arXiv ↗
14

VACE: All-in-One Video Creation and Editing

Zeyinzi Jiang, Zhen Han, Chaojie Mao +3 authors

VACE, an all-in-one framework for video creation and editing, integrates multiple tasks within a unified model using a Video Condition Unit and Context Adapter for flexible and consistent video synthesis.

58diffusion transformervideo synthesisHF ↗arXiv ↗
19

Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Yuxiao Qu, Matthew Y. R. Yang, Amrith Setlur +4 authors

The paper formalizes test-time compute optimization as a meta-reinforcement learning problem, introducing Meta Reinforcement Fine-Tuning (MRT) to enhance performance and token efficiency in large language models.

48meta-reinforcement learningRLHF ↗arXiv ↗
20

S2S-Arena, Evaluating Speech2Speech Protocols on Instruction Following with Paralinguistic Information

Feng Jiang, Zhiyu Lin, Fan Bu +3 authors

S2S-Arena is introduced to evaluate speech models' instruction-following abilities with paralinguistic information, revealing that cascaded ASR, LLM, and TTS outperform jointly trained models in speech2speech protocols and that generating appropriate audio with paralinguistic information remains challenging.

47large language modelsLLMsHF ↗arXiv ↗
22

TPDiff: Temporal Pyramid Video Diffusion Model

Lingmin Ran, Mike Zheng Shou

A multi-stage diffusion framework, TPDiff, enhances video diffusion model efficiency by reducing full frame rate during high-entropy stages, leading to diminished training costs and improved inference efficiency.

45video diffusion modelsentropy-reducingHF ↗arXiv ↗
23

Automated Movie Generation via Multi-Agent CoT Planning

Weijia Wu, Zeyu Zhu, Mike Zheng Shou

MovieAgent automates movie generation through a hierarchical Chain of Thought framework using multiple language models to handle narrative structuring, ensuring script fidelity, character consistency, and narrative coherence.

44Chain of Thought (CoT)multi-agent systemHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号