TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Mar 2 – Mar 8, 2026

50 篇论文 · 按点赞排序

31

Kling-MotionControl Technical Report

Kling Team, Jialu Chen, Yikang Ding +21 authors

Kling-MotionControl is a DiT-based framework for character animation that combines heterogeneous motion representations, adaptive identity-agnostic learning, and advanced acceleration techniques to achieve high-fidelity, expressive, and controllable video generation.

27DiT-based frameworkheterogeneous motion representationsHF ↗arXiv ↗
32

Large Multimodal Models as General In-Context Classifiers

Marco Garosi, Matteo Farina, Alessandro Conti +2 authors

Large Multimodal Models demonstrate superior performance in closed-world classification with in-context learning and excel in open-world scenarios when equipped with context refinement techniques.

27Vision-Language ModelsLarge Multimodal ModelsHF ↗arXiv ↗
38

LLaDA-o: An Effective and Length-Adaptive Omni Diffusion Model

Zebin You, Xiaolu Zhang, Jun Zhou +2 authors

LLaDA-o is an omni diffusion model that uses a Mixture of Diffusion framework to jointly handle text understanding and visual generation through a shared attention backbone, achieving state-of-the-art performance in multimodal tasks.

22Mixture of Diffusionomni diffusion modelHF ↗arXiv ↗
40

Phi-4-reasoning-vision-15B Technical Report

Jyoti Aneja, Michael Harrison, Neel Joshi +3 authors

A compact open-weight multimodal reasoning model is presented that achieves competitive performance through careful architecture design, high-quality data curation, and a hybrid approach combining direct answering with chain-of-thought reasoning.

21multimodal reasoning modelopen-weight modelHF ↗arXiv ↗
41

Next Embedding Prediction Makes World Models Stronger

George Bredis, Nikita Balagansky, Daniil Gavrilov +1 authors

NE-Dreamer uses a temporal transformer to predict next-step encoder embeddings for model-based reinforcement learning without requiring decoders or auxiliary supervision.

21temporal transformernext-step encoder embeddingsHF ↗arXiv ↗
43

Interactive Benchmarks

Baoqing Yue, Zihan Zhu, Yifan Zhang +3 authors

Interactive benchmarks offer a unified framework for evaluating model intelligence through active information acquisition under constraint conditions, demonstrating superior assessment of reasoning capabilities compared to traditional benchmarks.

19interactive benchmarksmodel intelligenceHF ↗arXiv ↗
44

Memex(RL): Scaling Long-Horizon LLM Agents via Indexed Experience Memory

Zhenting Wang, Huancheng Chen, Jiayun Wang +1 authors

A memory mechanism called Memex enables large language model agents to handle long-horizon tasks more effectively by maintaining compact context through structured summaries while storing full interaction details in an external database, allowing selective retrieval based on learned criteria.

19large language model agentscontext windowsHF ↗arXiv ↗
45

SageBwd: A Trainable Low-bit Attention

Jintao Zhang, Marco Chen, Haoxu Wang +5 authors

Research investigates why low-bit attention methods like SageBwd exhibit performance gaps during pre-training and identifies key factors for stable training with reduced precision.

19SageAttentionINT8 attentionHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号