TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

686 篇论文 · 按点赞排序

361

OmniScientist: An Omni-Modal Omni-Discipline AI Scientist

Bobo Li, Hao Fei, Tianjie Ju +2 authors

OmniScientist is an end-to-end omni-modal AI scientist that performs multidisciplinary research directly from heterogeneous raw evidence using autonomous agents and lifecycle-wide perception, improving evidence-grounded discovery across diverse scientific modalities.

91omni-modal AI scientistperception layerHF ↗arXiv ↗
362

AOrchestra: Automating Sub-Agent Creation for Agentic Orchestration

Jianhao Ruan, Zhihao Xu, Yiran Peng +8 authors

AOrchestra is a framework-agnostic agentic system that uses a tuple-based abstraction to dynamically create specialized task executors, achieving improved performance on complex benchmarks through automated agent creation and resource management.

90language agentssub-agent-as-tools paradigmHF ↗arXiv ↗
366

Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning

Chen Tang, Yizhou Wang, Jianyu Wu +26 authors

SciReasoner is a multimodal scientific foundation model that enables interpretable structural reasoning across proteins, molecules, and crystals by discretizing structural elements into a unified vocabulary for enhanced prediction and scientific inference.

89multimodal scientific foundation modelstructural reasoningHF ↗arXiv ↗
368

DFlash: Block Diffusion for Flash Speculative Decoding

Jian Chen, Yesheng Liang, Zhijian Liu

DFlash is a speculative decoding framework that uses a lightweight block diffusion model for parallel token drafting, achieving significant speedup over existing autoregressive methods while maintaining high-quality outputs.

89autoregressive large language modelsspeculative decodingHF ↗arXiv ↗
370

UniG2U-Bench: Do Unified Models Advance Multimodal Understanding?

Zimo Wen, Boxiu Li, Wanbo Zhang +11 authors

Unified multimodal models show mixed performance in generation-to-understanding tasks, with specific subtasks benefiting from enhanced spatial and reasoning capabilities while overall performance lags behind specialized vision-language models.

88Unified multimodal modelsVision-Language ModelsHF ↗arXiv ↗
371

Evolving Programmatic Skill Networks

Haochen Shi, Xingdi Yuan, Bang Liu

Programmatic Skill Network enables continual skill acquisition through executable symbolic programs that evolve via reflection, progressive optimization, and structural refactoring mechanisms.

88Programmatic Skill Networkexecutable symbolic programsHF ↗arXiv ↗
373

Video Generation Models are General-Purpose Vision Learners

Letian Wang, Chuhan Zhang, Rishabh Kabra +9 authors

GenCeption uses a video generative diffusion backbone to build a generalist vision model that achieves strong performance across diverse perception tasks with high data efficiency and emergent generalization.

88text-to-video generationgenerative diffusionHF ↗arXiv ↗
379

GEMS: Agent-Native Multimodal Generation with Memory and Skills

Zefeng He, Siyuan Huang, Xiaoye Qu +4 authors

GEMS is an agent-native multimodal generation framework that enhances model capabilities through structured multi-agent optimization, persistent memory, and domain-specific skills across general and downstream tasks.

87multimodal generation modelsagent frameworksHF ↗arXiv ↗
380

LLM-in-Sandbox Elicits General Agentic Intelligence

Daixuan Cheng, Shaohan Huang, Yuxian Gu +6 authors

LLM-in-Sandbox enables large language models to perform general intelligence tasks across diverse domains by allowing them to explore a code sandbox environment, achieving robust generalization without additional training.

87LLM-in-Sandboxcode sandboxHF ↗arXiv ↗
382

OpenGame: Open Agentic Coding for Games

Yilei Jiang, Jinyuan Hu, Qianyin Xiao +8 authors

OpenGame is an open-source agentic framework for end-to-end web game creation that uses specialized code models and evaluation benchmarks to overcome challenges in interactive application development.

86Large Language Modelscode agentsHF ↗arXiv ↗
384

Programmable World Model

Zheng-Hui Huang, Guixu Lin, Jiacheng Lin +8 authors

A programmable world model separates explicit state evolution from video generation using executable rules and 3D bounding boxes to maintain persistent, controllable environments.

86programmable world modelexecutable programsHF ↗arXiv ↗
13 / 23

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号