TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Mar 9 – Mar 15, 2026

50 篇论文 · 按点赞排序

34

Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs

Kaiser Sun, Xiaochuang Yuan, Hongjun Liu +4 authors

Multimodal large language models exhibit inconsistent performance when processing text from images versus textual tokens, with factors like rendering quality and task type influencing this modality gap, which can be mitigated through self-distillation techniques that leverage text-based reasoning traces.

29multimodal large language modelsmodality gapHF ↗arXiv ↗
35

DVD: Deterministic Video Depth Estimation with Generative Priors

Hongfei Zhang, Harold Haodong Chen, Chenfei Liao +12 authors

Video depth estimation framework DVD adapts pre-trained video diffusion models into deterministic single-pass depth regressors using structural anchors, latent manifold rectification, and global affine coherence for improved accuracy and efficiency.

28video diffusion modelsdepth regressorsHF ↗arXiv ↗
38

ReMix: Reinforcement routing for mixtures of LoRAs in LLM finetuning

Ruizhong Qiu, Hanqing Zeng, Yinglong Xia +15 authors

Researchers address imbalance in routing weights of Mixture-of-LoRAs models by proposing Reinforcement Routing (ReMix), which uses non-learnable weights and reinforcement learning techniques to improve model expressiveness and performance.

26low-rank adaptersparameter-efficient fine-tuningHF ↗arXiv ↗
42

NLE: Non-autoregressive LLM-based ASR by Transcript Editing

Avihu Dekel, Samuel Thomas, Takashi Fukada +1 authors

A non-autoregressive speech recognition approach formulates acoustic-to-text conversion as conditional transcript editing using a bidirectional language model editor with latent alignment training and interleaved padding for improved efficiency.

24autoregressivenon-autoregressiveHF ↗arXiv ↗
44

One Model, Many Budgets: Elastic Latent Interfaces for Diffusion Transformers

Moayed Haji-Ali, Willi Menapace, Ivan Skorokhodov +6 authors

Elastic Latent Interface Transformer (ELIT) decouples compute from image resolution in diffusion transformers by introducing learnable latent tokens that adaptively prioritize important regions, enabling dynamic resource allocation without modifying core model architecture.

18diffusion transformersFLOPsHF ↗arXiv ↗
45

V_{0.5}: Generalist Value Model as a Prior for Sparse RL Rollouts

Yi-Kai Zhang, Yueqing Sun, Hongyan Hao +4 authors

Adaptive value estimation method combines pretrained prior with empirical rollouts using real-time statistical testing to reduce variance and improve reinforcement learning performance under sparse sampling conditions.

17policy gradientsGeneralist Value ModelsHF ↗arXiv ↗
46

RubiCap: Rubric-Guided Reinforcement Learning for Dense Image Captioning

Tzu-Heng Huang, Sirajul Salekin, Javier Movellan +2 authors

RubiCap introduces a reinforcement learning framework for dense image captioning that uses large language model-generated rubrics to provide fine-grained reward signals, achieving superior performance over supervised methods and competing with larger models.

17dense image captioningvision-language modelsHF ↗arXiv ↗
49

Scale Space Diffusion

Soumik Mukhopadhyay, Prateksha Udhayanan, Abhinav Shrivastava

Scale-space theory connects diffusion models' information hierarchy to low-pass filtering, leading to a framework that combines scale spaces with diffusion processes for efficient image processing.

16diffusion modelsscale-space theoryHF ↗arXiv ↗
50

Dynamic Chunking Diffusion Transformer

Akash Haridas, Utkarsh Saxena, Parsa Ashrafi Fashi +3 authors

Dynamic Chunking Diffusion Transformer adapts token sequence length based on image content and diffusion timestep, improving efficiency and performance over fixed-token approaches.

16Diffusion TransformersDiTHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号