TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Mar 31 – Apr 6, 2025
本周最热305

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems

Bang Liu, Xinfeng Li, Jiayi Zhang +44 authors

This survey covers the design, evaluation, and improvement of intelligent agents based on modular, brain-inspired architectures, focusing on self-enhancement, multi-agent collaboration, and safety in AI systems.

large language modelsintelligent agentsmodular architecturebrain-inspiredHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

MoCha: Towards Movie-Grade Talking Character Synthesis

Cong Wei, Bo Sun, Haoyu Ma +10 authors

MoCha generates realistic talking character animations from speech and text using a speech-video attention mechanism and joint training on speech-labeled and text-labeled data, enabling multi-character conversations and superior realism.

141MoChaspeech-video window attention mechanismHF ↗arXiv ↗
05

ZClip: Adaptive Spike Mitigation for LLM Pre-Training

Abhay Kumar, Louis Owen, Nilabhra Roy Chowdhury +1 authors

ZClip is an adaptive gradient clipping algorithm that uses z-score-based anomaly detection to dynamically adjust clipping thresholds and prevent large gradient spikes during LLM training.

90gradient instabilityloss spikesHF ↗arXiv ↗
10

Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Zhenyi Liao, Qingsong Xie, Yanhao Zhang +4 authors

The study enhances visual-spatial reasoning in multi-modal large language models through GRPO training using the VSI-100k dataset, demonstrating significant performance improvements over base models.

67multi-modal large language modelsvisual-spatial intelligenceHF ↗arXiv ↗
17

Multi-Token Attention

Olga Golovneva, Tianlu Wang, Jason Weston +1 authors

A new attention mechanism, Multi-Token Attention (MTA), enhances LLM performance by conditioning attention weights on multiple query and key vectors, improving context search and standard language modeling tasks.

56soft attentionLLMsHF ↗arXiv ↗
21

Efficient Inference for Large Reasoning Models: A Survey

Yue Liu, Jiaying Wu, Yufei He +6 authors

This survey reviews efficient inference methods for Large Reasoning Models to reduce token usage and memory consumption while maintaining reasoning quality, categorizing methods into explicit compact Chain-of-Thought and implicit latent CoT.

45Large Reasoning ModelsLarge Language ModelsHF ↗arXiv ↗
28

SkyReels-A2: Compose Anything in Video Diffusion Transformers

Zhengcong Fei, Debang Li, Di Qiu +8 authors

SkyReels-A2, an open-source framework, generates high-quality, element-controlled videos from textual prompts using a novel image-text embedding model, optimized inference pipeline, and A2 Bench for systematic evaluation.

39elements-to-video (E2V)image-text joint embeddingHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号