TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

606 篇论文 · 按点赞排序

424

Continuous Autoregressive Language Models

Chenze Shao, Darren Li, Fandong Meng +1 authors

Continuous Autoregressive Language Models (CALM) improve language model efficiency by predicting continuous vectors instead of discrete tokens, reducing computational cost while maintaining performance.

74Continuous Autoregressive Language ModelsCALMHF ↗arXiv ↗
426

YuE: Scaling Open Foundation Models for Long-Form Music Generation

Ruibin Yuan, Hanfeng Lin, Shuyue Guo +54 authors

YuE, a family of open foundation models based on LLaMA2, can generate long-form music with aligned lyrics, coherent structure, and appropriate accompaniment using innovative techniques in next-token prediction, conditioning, and pre-training.

73track-decoupled next-token predictionstructural progressive conditioningHF ↗arXiv ↗
432

Deep Research: A Systematic Survey

Zhengliang Shi, Yiqun Chen, Haitao Li +23 authors

Deep Research systems integrate LLMs with external tools to enhance problem-solving capabilities, involving query planning, information acquisition, memory management, and answer generation.

73Deep ResearchLarge language modelsHF ↗arXiv ↗
433

Thus Spake Long-Context Large Language Model

Xiaoran Liu, Ruixiao Li, Mianqiu Huang +10 authors

The survey examines the advancements and challenges in long-context Large Language Models (LLMs), exploring the architecture, infrastructure, training, and evaluation technologies needed to extend their context length and address the inherent trade-offs.

73Large Language Modelslong contextHF ↗arXiv ↗
434

RewardDance: Reward Scaling in Visual Generation

Jie Wu, Yu Gao, Zilyu Ye +9 authors

RewardDance is a scalable reward modeling framework that aligns with VLM architectures, enabling effective scaling of RMs and resolving reward hacking issues in generation models.

73CLIP-based RMsBradley-Terry lossesHF ↗arXiv ↗
439

Matrix-Game: Interactive World Foundation Model

Yifan Zhang, Chunli Peng, Boyang Wang +8 authors

Matrix-Game, a controllable game world generation model trained in a two-stage process, outperforms existing models by producing high-quality, action-controllable, and physically consistent Minecraft world videos.

72Matrix-Gameinteractive world foundation modelHF ↗arXiv ↗
442

System Prompt Optimization with Meta-Learning

Yumin Choi, Jinheon Baek, Sung Ju Hwang

A meta-learning framework for optimizing system prompts in Large Language Models (LLMs) improves generalization across diverse tasks and datasets.

72Large Language Models (LLMs)bilevel system prompt optimizationHF ↗arXiv ↗
443

Qwen2.5-1M Technical Report

An Yang, Bowen Yu, Chengyuan Li +25 authors

The Qwen2.5-1M series models extend context length to 1 million tokens with enhanced long-context capabilities, employing techniques like long data synthesis and progressive pre-training, and are supported by an open-source inference framework with sparse attention and kernel optimizations.

72long-context pre-traininglong data synthesisHF ↗arXiv ↗
444

Enabling Scalable Oversight via Self-Evolving Critic

Zhengyang Tang, Ziniu Li, Zhenyang Xiao +8 authors

SCRIT, a self-evolving critique framework, enhances LLMs' critique capabilities using synthetic data and self-validation, achieving significant improvements in critique-correction and error identification benchmarks.

72Large Language Models (LLMs)self-evolvingHF ↗arXiv ↗
447

Part-X-MLLM: Part-aware 3D Multimodal Large Language Model

Chunshi Wang, Junliang Ye, Yunhan Yang +6 authors

We introduce Part-X-MLLM, a native 3D multimodal large language model that unifies diverse 3D tasks by formulating them as programs in a structured, executable grammar. Given an RGB point cloud and a natural language prompt, our model autoregressively generates a single, coherent token sequence encoding part-level bounding boxes, semantic descriptions, and edit commands. This structured output serves as a versatile interface to drive downstream geometry-aware modules for part-based generation and editing. By decoupling the symbolic planning from the geometric synthesis, our approach allows any compatible geometry engine to be controlled through a single, language-native frontend. We pre-train a dual-encoder architecture to disentangle structure from semantics and instruction-tune the model on a large-scale, part-centric dataset. Experiments demonstrate that our model excels at producing high-quality, structured plans, enabling state-of-the-art performance in grounded Q\&A, compositional generation, and localized editing through one unified interface. Project page: https://chunshi.wang/Part-X-MLLM/

72HF ↗arXiv ↗
449

Energy-Based Transformers are Scalable Learners and Thinkers

Alexi Gladstone, Ganesh Nanduru, Md Mofijul Islam +7 authors

Energy-Based Transformers (EBTs) improve model performance and scalability across modalities by learning to verify predictions through unsupervised learning and energy minimization.

71Energy-Based TransformersEnergy-Based ModelsHF ↗arXiv ↗
15 / 21

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号