TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Jun 1 – Jun 7, 2026
本周最热254

Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs

Haozhe Zhao, Shuzheng Si, Zhenhailong Wang +6 authors

Automated systems for generating scientific figures face limitations in handling diverse figure types and conditions, prompting the development of multi-agent frameworks that generalize across different input scenarios and produce editable output formats.

multi-agent harnessfigure generationraster outputseditable SVGsHF ↗arXiv ↗

50 篇论文 · 按点赞排序

03

Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding

Jianuo Huang, Yaojie Zhang, Qituan Zhang +3 authors

Domino is a speculative decoding framework that improves LLM inference speed by decoupling causal dependency modeling from autoregressive drafting through a parallel backbone and lightweight causal refinement head, achieving significant speedups in both end-to-end execution and throughput.

152speculative decodingautoregressive draftersHF ↗arXiv ↗
04

Cosmos 3: Omnimodal World Models for Physical AI

Aditi, Niket Agarwal, Arslan Ali +288 authors

Cosmos 3 is an omnimodal world model that processes and generates multiple data types through a unified mixture-of-transformers architecture, achieving state-of-the-art performance in various understanding and generation tasks.

143omnimodal world modelsmixture-of-transformers architectureHF ↗arXiv ↗
06

Audio Interaction Model

Zhifei Xie, Zihang Liu, Ze An +8 authors

A unified streaming audio model is developed that combines offline task execution with real-time audio instruction following through an end-to-end framework supporting multiple audio interaction capabilities.

124Large Audio Language Modelsstreaming audio modelsHF ↗arXiv ↗
08

OCC-RAG: Optimal Cognitive Core for Faithful Question Answering

Maksim Savkin, Mikhail Goncharov, Alexander Gambashidze +7 authors

Compact task-specialized language models demonstrate superior performance in multi-hop reasoning and faithfulness compared to larger general-purpose models through a novel training pipeline and structured reasoning traces.

103language modelstask-specialized modelsHF ↗arXiv ↗
12

Trust-Region Behavior Blending for On-Policy Distillation

Daniil Plyusov, Alexey Gorbatovski, Alexey Malakhov +4 authors

Trust-Region behavior Blending improves on-policy distillation by replacing early poor-quality student rollouts with teacher-like behavior within a KL trust region during warmup.

69on-policy distillationstudent policyHF ↗arXiv ↗
15

Representation Forcing for Bottleneck-Free Unified Multimodal Models

Yuqing Wang, Zhijie Lin, Ceyuan Yang +10 authors

Representation Forcing enables unified multimodal models to perform both perception and generation tasks end-to-end without relying on external latent spaces, matching state-of-the-art performance in image generation while improving understanding capabilities.

64unified multimodal modelsvisual representationsHF ↗arXiv ↗
18

Mellum2 Technical Report

Marko Kojic, Ivan Bondyrev, Aral de Moor +6 authors

Mellum 2 is an open-weight 12B-parameter Mixture-of-Experts language model with 2.5B active parameters per token, specialized in software engineering tasks and optimized for inference efficiency on commodity GPUs.

60Mixture-of-ExpertsGrouped-Query AttentionHF ↗arXiv ↗
21

Benchmarking Visual State Tracking in Multimodal Video Understanding

Sihyun Yu, Nanye Ma, Pinzhi Huang +8 authors

Current multimodal large language models struggle with visual state tracking in videos, performing poorly even when human-level capabilities are required, and existing agentic approaches do not effectively address these limitations.

54Multimodal Large Language Modelsvisual state trackingHF ↗arXiv ↗
22

ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time?

Woojung Song, Nalim Kim, Sangjun Song +3 authors

Role-playing language agents require dynamic character development that evolves through narratives, necessitating benchmarks that evaluate psychological trajectory alignment rather than static factual recall, with ArcANE demonstrating superior performance when character arc information is conditioned into models.

50role-playing language agentscharacter arcHF ↗arXiv ↗
23

Trust Region On-Policy Distillation

Xingrun Xing, Haoqing Wang, Boyan Gao +2 authors

Trust Region On-Policy Distillation (TrOPD) improves reliable token-level supervision in large language model distillation by using trust regions, outlier estimation, and off-policy guidance to address instability issues under distribution mismatch.

48on-policy distillationtrust regionHF ↗arXiv ↗
30

Function2Scene: 3D Indoor Scene Layout from Functional Specifications

Ruiqi Wang, Qimin Chen, Daniel Ritchie +4 authors

Function2Scene generates 3D indoor layouts from functional descriptions by parsing user needs and applying design constraints through an iterative refinement process combining geometric analysis, language modeling, and visual assessment.

43text-driven 3D indoor scene synthesisfunctional specificationsHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号