TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

686 篇论文 · 按点赞排序

301

Generative World Renderer

Zheng-Hui Huang, Zhixiang Wang, Jiaming Tan +6 authors

A large-scale dynamic dataset derived from AAA games is introduced to improve generative inverse and forward rendering, featuring high-resolution synchronized RGB and G-buffer data alongside a novel VLM-based evaluation method that correlates well with human judgment.

103G-bufferinverse renderingHF ↗arXiv ↗
305

OCC-RAG: Optimal Cognitive Core for Faithful Question Answering

Maksim Savkin, Mikhail Goncharov, Alexander Gambashidze +7 authors

Compact task-specialized language models demonstrate superior performance in multi-hop reasoning and faithfulness compared to larger general-purpose models through a novel training pipeline and structured reasoning traces.

102language modelstask-specialized modelsHF ↗arXiv ↗
307

Flow-OPD: On-Policy Distillation for Flow Matching Models

Zhen Fang, Wenxuan Huang, Yu Zeng +8 authors

Flow-OPD addresses limitations in Flow Matching text-to-image models through a two-stage alignment approach combining on-policy distillation and manifold anchor regularization, achieving significant improvements in generation quality and alignment metrics.

102Flow Matchingon-policy distillationHF ↗arXiv ↗
309

ABot-N1: Toward a General Visual Language Navigation Foundation Model

Ruiyan Gong, Yingnan Guo, Junjun Hu +37 authors

ABot-N1 improves visual language navigation by separating reasoning from control through a slow-fast architecture that uses explicit chain-of-thought reasoning and pixel goals to guide continuous waypoint generation, achieving strong results across diverse embodied tasks.

102Visual Language NavigationChain-of-Thought reasoningHF ↗arXiv ↗
315

DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution

Jiashu Zhu, Yanhao Zheng, Ruitian Tian +7 authors

A compact 7B native joint audio-video generator uses cross-modal attention, progressive joint training, reinforcement learning with multimodal feedback, and an autoregressive 2K refinement pipeline to produce synchronized high-resolution outputs.

100Gated Cross-Modal Attentiontoken- and head-wise output gatesHF ↗arXiv ↗
319

Active Learners as Efficient PRP Rerankers

Jeremías Figueiredo Paschmann, Juan Kaplan, Francisco Nattero +3 authors

Pairwise ranking prompting is reformulated as active learning from noisy comparisons, with improved rankers that enhance ranking quality under call constraints and address position bias through a randomized oracle.

98pairwise ranking promptingactive learningHF ↗arXiv ↗
320

Terminal Agents Suffice for Enterprise Automation

Patrice Bechard, Orlando Marquez Ayala, Emily Chen +5 authors

Simple terminal-based coding agents using programmatic interfaces and foundation models can effectively perform enterprise tasks comparable to or better than complex tool-augmented agents.

98tool-augmented agentsModel Context ProtocolHF ↗arXiv ↗
324

FORGE:Fine-grained Multimodal Evaluation for Manufacturing Scenarios

Xiangru Jian, Hao Xu, Wei Pang +13 authors

FORGE introduces a high-quality multimodal manufacturing dataset with fine-grained domain semantics to evaluate MLLMs on real-world tasks, revealing that domain-specific knowledge rather than visual grounding limits performance, and demonstrating that supervised fine-tuning on structured annotations significantly improves accuracy.

97Multimodal Large Language Modelsvisual groundingHF ↗arXiv ↗
329

MAXS: Meta-Adaptive Exploration with LLM Agents

Jian Zhang, Zhiyuan Wang, Zhangqi Wang +7 authors

MAXS is a meta-adaptive reasoning framework for LLM agents that improves multi-tool reasoning through lookahead strategies and trajectory convergence mechanisms, balancing global effectiveness and computational efficiency.

96LLM agentstool executionHF ↗arXiv ↗
11 / 23

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号