TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Dec 16 – Dec 22, 2024
本周最热381

Qwen2.5 Technical Report

Qwen, An Yang, Baosong Yang +39 authors

Qwen2.5, an enhanced series of large language models, demonstrates superior performance across various benchmarks and use cases through extensive pre-training and advanced post-training techniques.

large language modelspre-trainingpost-trainingsupervised finetuningHF ↗arXiv ↗

50 篇论文 · 按点赞排序

04

Byte Latent Transformer: Patches Scale Better Than Tokens

Artidoro Pagnoni, Ram Pasunuru, Pedro Rodriguez +11 authors

A Byte Latent Transformer (BLT) matches tokenization-based LLM performance at scale with improved inference efficiency and robustness, using entropy-based byte patching.

104Byte Latent Transformer (BLT)byte-level LLMHF ↗arXiv ↗
05

GenEx: Generating an Explorable World

Taiming Lu, Tianmin Shu, Junfei Xiao +8 authors

GenEx generates 3D environments from a single image, enabling AI agents to explore and interact with a consistent, expansive space through guided generative imagination.

98Generative imaginationpanoramic video streamsHF ↗arXiv ↗
08

Progressive Multimodal Reasoning via Active Retrieval

Guanting Dong, Chenghao Zhang, Mengjie Deng +3 authors

AR-MCTS enhances multimodal large language models' reasoning capabilities through active retrieval, Monte Carlo Tree Search, and a process reward model, improving performance across multimodal reasoning tasks.

73Active RetrievalMonte Carlo Tree SearchHF ↗arXiv ↗
10

AniDoc: Animation Creation Made Easier

Yihao Meng, Hao Ouyang, Hanlin Wang +6 authors

AniDoc uses video diffusion models to automate colorization and in-betweening in 2D animation, improving efficiency by leveraging correspondence matching.

58video diffusion modelscorrespondence matchingHF ↗arXiv ↗
11

How to Synthesize Text Data without Model Collapse?

Xuekai Zhu, Daixuan Cheng, Hengli Li +7 authors

The use of synthetic data in language model training leads to model collapse, which is mitigated by token-level editing of human-produced data to create semi-synthetic data.

52synthetic datamodel collapseHF ↗arXiv ↗
17

BrushEdit: All-In-One Image Inpainting and Editing

Yaowei Li, Yuxuan Bian, Xuan Ju +3 authors

BrushEdit addresses image editing limitations by combining multimodal large language models and dual-branch inpainting models for autonomous and interactive free-form instruction editing.

36diffusion modelsMLLMsHF ↗arXiv ↗
20

Large Action Models: From Inception to Implementation

Lu Wang, Fangkai Yang, Chaoyun Zhang +15 authors

A framework is proposed for developing Large Action Models (LAMs) that generate and execute actions in dynamic environments, moving beyond traditional Large Language Models (LLMs) towards intelligent agents capable of task completion.

36Large Action ModelsLAMsHF ↗arXiv ↗
24

GUI Agents: A Survey

Dang Nguyen, Jian Chen, Yu Wang +26 authors

A survey of graphical user interface agents powered by large foundation models, detailing benchmarks, metrics, architectures, training methods, and future challenges in automating human-computer interaction.

29Large Foundation ModelsGUI agentsHF ↗arXiv ↗
25

Smaller Language Models Are Better Instruction Evolvers

Tingfeng Hui, Lulu Zhao, Guanting Dong +3 authors

Smaller language models can generate more effective and diverse instructions than larger models, challenging the assumption that model size directly correlates with instruction synthesis quality.

29instruction tuninglarge language modelsHF ↗arXiv ↗
27

ColorFlow: Retrieval-Augmented Image Sequence Colorization

Junhao Zhuang, Xuan Ju, Zhaoyang Zhang +4 authors

ColorFlow, a three-stage diffusion-based framework with a dual-branch design, achieves superior image sequence colorization by incorporating retrieval augmented colorization and self-attention mechanisms, outperforming existing models.

25diffusion-based frameworkretrieval augmented colorizationHF ↗arXiv ↗
29

Causal Diffusion Transformers for Generative Modeling

Chaorui Deng, Deyao Zh, Kunchang Li +2 authors

Causal Diffusion combines autoregressive models with diffusion techniques to improve generation performance and multimodal capabilities, achieving state-of-the-art results in image generation and zero-shot image manipulations.

23Causal Diffusionautoregressive (AR)HF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号