TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Dec 2 – Dec 8, 2024
本周最热134

PaliGemma 2: A Family of Versatile VLMs for Transfer

Andreas Steiner, André Susano Pinto, Michael Tschannen +15 authors

PaliGemma 2 integrates a SigLIP-So400m vision encoder with Gemma 2 models of varying sizes and resolutions, advancing transfer performance across diverse vision-language tasks, including OCR and captioning.

SigLIP-So400mvision encoderGemma 2 modelstransfer performanceHF ↗arXiv ↗

50 篇论文 · 按点赞排序

09

NVILA: Efficient Frontier Visual Language Models

Zhijian Liu, Ligeng Zhu, Baifeng Shi +24 authors

NVILA, a family of VLMs, optimizes efficiency and accuracy through a scale-then-compress approach, enhancing performance across various benchmarks while reducing computational costs.

61Visual language modelsNVILAHF ↗arXiv ↗
13

MALT: Improving Reasoning with Multi-Agent LLM Training

Sumeet Ramesh Motwani, Chandler Smith, Rocktim Jyoti Das +6 authors

Multi-agent LLM training improves performance on reasoning tasks by assigning specialized roles and utilizing joint outcome-based rewards to enhance collaboration among models.

46sequential multi-agent setupheterogeneous LLMsHF ↗arXiv ↗
14

GRAPE: Generalizing Robot Policy via Preference Alignment

Zijian Zhang, Kaiyuan Zheng, Zhaorun Chen +6 authors

GRAPE improves vision-language-action models' performance by aligning policies via preference modeling, enhancing generalizability and allowing customization of objectives like safety and efficiency.

46vision-language-action modelsGRAPEHF ↗arXiv ↗
15

o1-Coder: an o1 Replication for Coding

Yuxiang Zhang, Shangxi Wu, Yuqi Yang +4 authors

O1-CODER integrates reinforcement learning and Monte Carlo Tree Search to enhance coding capabilities, focusing on pseudocode and full code generation through iterative fine-tuning and standardized code testing.

45reinforcement learningMonte Carlo Tree SearchHF ↗arXiv ↗
16

Video Depth without Video Models

Bingxin Ke, Dominik Narnhofer, Shengyu Huang +5 authors

RollingDepth transforms a single-image latent diffusion model into a video depth estimator by combining multi-frame depth estimation with optimization-based registration for consistent long-video depth prediction.

38latent diffusion modelvideo depth estimationHF ↗arXiv ↗
20

Free Process Rewards without Process Labels

Lifan Yuan, Wendi Li, Huayu Chen +6 authors

An implicit process reward model can be trained without additional step labels by parameterizing outcome rewards as log-likelihood ratios, outperforming traditional methods with less data and better generalization.

34process reward modelPRMHF ↗arXiv ↗
21

Switti: Designing Scale-Wise Transformers for Text-to-Image Synthesis

Anton Voronov, Denis Kuznedelev, Mikhail Khoroshikh +2 authors

Switti, a scale-wise transformer for text-to-image generation, improves convergence and performance through architectural modifications, reduces memory usage, and achieves competitive results compared to diffusion models with significant speed advantages.

34scale-wise transformertext-to-image generationHF ↗arXiv ↗
22

Open-Sora Plan: Open-Source Large Video Generation Model

Bin Lin, Yunyang Ge, Xinhua Cheng +21 authors

Open-Sora Plan is an open-source project that generates high-resolution, long-duration videos from various inputs using a Wavelet-Flow Variational Autoencoder, Joint Image-Video Skiparse Denoiser, and condition controllers.

33Wavelet-Flow Variational AutoencoderJoint Image-Video Skiparse DenoiserHF ↗arXiv ↗
24

On Domain-Specific Post-Training for Multimodal Large Language Models

Daixuan Cheng, Shaohan Huang, Ziyu Zhu +5 authors

The paper explores the domain adaptation of multimodal large language models through post-training, utilizing a visual instruction synthesizer and single-stage training pipeline to improve performance on specific tasks in domains like biomedicine and food.

30multimodal large language modelsdomain adaptationHF ↗arXiv ↗
25

A Noise is Worth Diffusion Guidance

Donghoon Ahn, Jiwon Kang, Sanghyun Lee +9 authors

A noise-refining method is introduced to eliminate the need for guidance in diffusion models, enhancing image generation speed and memory efficiency.

29diffusion modelsclassifier-free guidanceHF ↗arXiv ↗
27

Yi-Lightning Technical Report

01. AI, Alan Wake, Albert Wang +39 authors

Yi-Lightning, a large language model with enhanced Mixture-of-Experts architecture and RAISE safety framework, achieves top-tier performance in specialized categories while being cost-effective and addressing safety issues.

29large language model (LLM)Mixture-of-Experts (MoE)HF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号