TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

May 19 – May 25, 2025

50 篇论文 · 按点赞排序

33

Scaling Diffusion Transformers Efficiently via μP

Chenyu Zheng, Xinyu Zhang, Rongzhen Wang +5 authors

Maximal Update Parametrization (μP) is extended to diffusion Transformers, demonstrating efficient hyperparameter transferability and reduced tuning costs across various models and tasks.

34Diffusion TransformersMaximal Update ParametrizationHF ↗arXiv ↗
35

Visual Agentic Reinforcement Fine-Tuning

Ziyu Liu, Yuhang Zang, Yushan Zou +6 authors

Visual Agentic Reinforcement Fine-Tuning enhances Large Vision-Language Models for flexible image manipulation and web-based reasoning, outperforming existing models on multi-modal benchmarks.

32Visual Agentic Reinforcement Fine-TuningLarge Vision-Language ModelsHF ↗arXiv ↗
36

Latent Flow Transformer

Yen-Chen Wu, Feng-Ting Liao, Meng-Hsi Chen +3 authors

The Latent Flow Transformer (LFT) compresses layers by replacing them with learned transport operators using flow matching and Flow Walking, showing improved performance over layer skipping and reducing the gap between autoregressive and flow-based models.

29Latent Flow TransformerLFTHF ↗arXiv ↗
41

The Aloe Family Recipe for Open and Specialized Healthcare LLMs

Dario Garcia-Gasulla, Jordi Bayarri-Planas, Ashwin Kumar Gururajan +10 authors

Aloe Beta models improve open-source medical Large Language Models through enhanced preprocessing, Direct Preference Optimization, Retrieval-Augmented Generation, and a rigorous evaluation methodology, achieving competitive performance and high safety standards.

26Large Language ModelsLLMsHF ↗arXiv ↗
44

EfficientLLM: Efficiency in Large Language Models

Zhengqing Yuan, Weixiang Sun, Yixin Liu +13 authors

EfficientLLM evaluates efficiency techniques for LLMs across architecture pretraining, fine-tuning, and inference, demonstrating task-dependent optima and cross-modal generalization.

25LLMsEfficientLLMHF ↗arXiv ↗
46

Training-Free Efficient Video Generation via Dynamic Token Carving

Yuechen Zhang, Jinbo Xing, Bin Xia +6 authors

Jenga, a novel inference pipeline for video Diffusion Transformer models, combines dynamic attention carving and progressive resolution generation to significantly speed up video generation while maintaining high quality.

24video Diffusion Transformer (DiT)self-attentionHF ↗arXiv ↗
47

Understanding Generative AI Capabilities in Everyday Image Editing Tasks

Mohammad Reza Taesiri, Brandon Collins, Logan Bolton +4 authors

Analysis of 83k image editing requests reveals that AI editors, including GPT-4o, struggle with low-creativity tasks and precise editing, while performing better on open-ended tasks, and human and VLM judges differ in their preferences for AI versus human edits.

24GPT-4oGemini-2.0-FlashHF ↗arXiv ↗
48

Constructing a 3D Town from a Single Image

Kaizhi Zheng, Ruijian Zhang, Jing Gu +2 authors

A training-free framework named 3DTown generates realistic 3D scenes from a single top-down image using region-based generation and spatial-aware 3D inpainting techniques.

243D generative models3DTownHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号