TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Mar 30 – Apr 5, 2026
本周最热366

DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language Models

Hao Liang, Zhengyang Zhao, Meiyi Qiang +22 authors

DataFlex is a unified framework for dynamic data-centric training of large language models that supports sample selection, domain mixture adjustment, and sample reweighting while maintaining compatibility with standard training workflows and enabling efficient large-scale deployment.

data-centric traininglarge language modelssample selectiondomain mixture adjustmentHF ↗arXiv ↗

50 篇论文 · 按点赞排序

04

Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models

Kaijin Chen, Dingkang Liang, Xin Zhou +4 authors

Hybrid Memory enables video world models to maintain consistent tracking of dynamic subjects during occlusion by combining archival storage for static backgrounds with active tracking for moving objects, using a specialized architecture with tokenized memory and spatiotemporal retrieval mechanisms.

158video world modelshybrid memoryHF ↗arXiv ↗
06

The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook

Xinlei Yu, Zhangquan Chen, Yongbo He +34 authors

Latent space is emerging as a fundamental computational substrate for language-based models, offering advantages over explicit token-level approaches through continuous representation that mitigates linguistic redundancy and sequential inefficiency.

153latent spacelanguage-based modelsHF ↗arXiv ↗
07

LongCat-Next: Lexicalizing Modalities as Discrete Tokens

Meituan LongCat Team, Bin Xiao, Chao Wang +86 authors

Discrete Native Autoregressive framework enables unified multimodal processing by representing diverse modalities in a shared discrete space through a novel visual transformer architecture.

151Next-Token Predictionautoregressive modelingHF ↗arXiv ↗
08

TAPS: Task Aware Proposal Distributions for Speculative Sampling

Mohamad Zbib, Mohamad Bazzi, Ammar Mohanna +2 authors

Speculative decoding effectiveness depends on draft model training data alignment with downstream tasks, with specialized drafters performing better when combined through confidence-based routing rather than simple averaging.

148speculative decodingdraft modelHF ↗arXiv ↗
09

Generative World Renderer

Zheng-Hui Huang, Zhixiang Wang, Jiaming Tan +6 authors

A large-scale dynamic dataset derived from AAA games is introduced to improve generative inverse and forward rendering, featuring high-resolution synchronized RGB and G-buffer data alongside a novel VLM-based evaluation method that correlates well with human judgment.

103G-bufferinverse renderingHF ↗arXiv ↗
11

Terminal Agents Suffice for Enterprise Automation

Patrice Bechard, Orlando Marquez Ayala, Emily Chen +5 authors

Simple terminal-based coding agents using programmatic interfaces and foundation models can effectively perform enterprise tasks comparable to or better than complex tool-augmented agents.

100tool-augmented agentsModel Context ProtocolHF ↗arXiv ↗
12

Towards a Medical AI Scientist

Hongtao Wu, Boyun Zheng, Dingjie Song +5 authors

Medical AI Scientist represents the first autonomous research framework designed for clinical applications, enabling evidence-based hypothesis generation and manuscript drafting through clinician-engineer collaboration across three research modes.

94autonomous research frameworkclinical autonomous researchHF ↗arXiv ↗
13

GEMS: Agent-Native Multimodal Generation with Memory and Skills

Zefeng He, Siyuan Huang, Xiaoye Qu +4 authors

GEMS is an agent-native multimodal generation framework that enhances model capabilities through structured multi-agent optimization, persistent memory, and domain-specific skills across general and downstream tasks.

88multimodal generation modelsagent frameworksHF ↗arXiv ↗
19

Steerable Visual Representations

Jona Ruthardt, Manu Gaur, Deva Ramanan +2 authors

Steerable Visual Representations enable language-guided focus on specific image elements while maintaining representation quality through early fusion of text and visual features.

58Vision TransformersDINOv2HF ↗arXiv ↗
20

Gen-Searcher: Reinforcing Agentic Search for Image Generation

Kaituo Feng, Manyuan Zhang, Shuang Chen +7 authors

A search-augmented image generation agent is presented that performs multi-hop reasoning and search to collect textual knowledge and reference images for grounded generation, trained with supervised fine-tuning and agentic reinforcement learning with dual reward feedback.

58search-augmented image generationmulti-hop reasoningHF ↗arXiv ↗
21

VOID: Video Object and Interaction Deletion

Saman Motamed, William Harvey, Benjamin Klein +3 authors

VOID is a video object removal framework that uses vision-language models and video diffusion models to generate physically plausible scenes by leveraging causal reasoning and counterfactual reasoning.

57video object removalvideo diffusion modelHF ↗arXiv ↗
27

CutClaw: Agentic Hours-Long Video Editing via Music Synchronization

Shifang Zhao, Yihan Hu, Ying Shan +2 authors

CutClaw is an autonomous multi-agent framework that uses multimodal language models to automatically edit long video footage into rhythmic, narratively consistent short videos with synchronized audio and visual elements.

51multimodal language modelsmulti-agent frameworkHF ↗arXiv ↗
29

EpochX: Building the Infrastructure for an Emergent Agent Civilization

Huacan Wang, Chaofa Yuan, Xialie Zhuang +15 authors

General-purpose technologies reshape economies less by improving individual tools than by enabling new ways to organize production and coordination. We believe AI agents are approaching a similar inflection point: as foundation models make broad task execution and tool use increasingly accessible, the binding constraint shifts from raw capability to how work is delegated, verified, and rewarded at scale. We introduce EpochX, a credits-native marketplace infrastructure for human-agent production networks. EpochX treats humans and agents as peer participants who can post tasks or claim them. Claimed tasks can be decomposed into subtasks and executed through an explicit delivery workflow with verification and acceptance. Crucially, EpochX is designed so that each completed transaction can produce reusable ecosystem assets, including skills, workflows, execution traces, and distilled experience. These assets are stored with explicit dependency structure, enabling retrieval, composition, and cumulative improvement over time. EpochX also introduces a native credit mechanism to make participation economically viable under real compute costs. Credits lock task bounties, budget delegation, settle rewards upon acceptance, and compensate creators when verified assets are reused. By formalizing the end-to-end transaction model together with its asset and incentive layers, EpochX reframes agentic AI as an organizational design problem: building infrastructures where verifiable work leaves persistent, reusable artifacts, and where value flows support durable human-agent collaboration.

47HF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号