Nash Learning from Human Feedback
Rémi Munos, Michal Valko, Daniele Calandriello +14 authors
NLHF uses a preference model and mirror descent to fine-tune LLMs for text summarization, advancing alignment with human preferences.
Explore · 每周精选
发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。
50 篇论文 · 按点赞排序
Rémi Munos, Michal Valko, Daniele Calandriello +14 authors
NLHF uses a preference model and mirror descent to fine-tune LLMs for text summarization, advancing alignment with human preferences.
Chuan Guo, Yuxuan Mu, Muhammad Gohar Javed +2 authors
MoMask, a masked modeling framework using hierarchical quantization and bidirectional transformers, excels in text-to-motion generation and related tasks with high fidelity and competitive performance.
Yuheng Jiang, Zhehao Shen, Penghao Wang +5 authors
HiFi4G uses a Gaussian-based approach with 3D Gaussian representation and non-rigid tracking for efficient high-fidelity human performance rendering, offering significant compression and quality improvements.
Jinxin Zhou, Tianyu Ding, Tianyi Chen +4 authors
DREAM, a novel training framework for diffusion models, improves training alignment with sampling, achieving faster convergence and reduced sampling steps in image super-resolution.
Zhonghao Wang, Wei Wei, Yang Zhao +4 authors
HiFi Tuner enhances object appearance conservation in personalized image generation through mask guidance, parameter regularization, and step-wise subject representations, achieving state-of-the-art results on the DreamBooth dataset.
Gege Gao, Weiyang Liu, Anpei Chen +2 authors
GraphDreamer generates complex 3D scenes from text inputs by using scene graphs to disentangle object entities and relationships, leveraging pretrained text-to-image diffusion models.
Jiayi Guo, Xingqian Xu, Yifan Pu +6 authors
Smooth Diffusion addresses non-smooth latent spaces in diffusion models by enforcing consistent variation ratios, improving tasks like text-to-image generation, interpolation, inversion, and editing.
Ivona Najdenkoska, Animesh Sinha, Abhimanyu Dubey +3 authors
Context Diffusion enhances in-context image generation by separately encoding visual context and preserving query image structure, improving quality and fidelity across different scenarios.
Lisa Dunlap, Yuhui Zhang, Xiaohan Wang +5 authors
VisDiff automatically generates descriptions of differences between two sets of images, aiding in the analysis of datasets and model behaviors across various applications.
Ali Hatamizadeh, Jiaming Song, Guilin Liu +2 authors
A novel diffusion model using vision transformers with a hierarchical architecture and time-dependent self-attention achieves state-of-the-art performance in image generation.
Zheqing Zhu, Rodrigo de Salvo Braz, Jalaj Bhandari +12 authors
Pearl is a modular, production-ready RL agent software package addressing various challenges in reinforcement learning.
Shariq Farooq Bhat, Niloy J. Mitra, Peter Wonka
LooseControl enables generalized depth conditioning for diffusion-based image generation using scene boundaries and 3D box control, enhancing flexibility and ease of use for creating complex environments.
Hao Zhang, Hongyang Li, Feng Li +8 authors
A new dataset and benchmark for grounded visual chat improve LMMs performance by integrating segmentation and language models.
Xinyu Zhang, Sebastian Hofstätter, Patrick Lewis +2 authors
The study develops effective listwise rerankers independent of GPT models, surpassing GPT-3.5 and achieving near-GPT-4 performance, and highlights the need for high-quality listwise ranking datasets.
Shunyuan Zheng, Boyao Zhou, Ruizhi Shao +4 authors
A GPS-Gaussian approach synthesizes 2K-resolution views in real-time using Gaussian parameter maps and depth estimation, outperforming existing methods.
Khai Loong Aw, Syrielle Montariol, Badr AlKhamissi +2 authors
Instruction-tuning improves brain alignment in large language models but not behavioral alignment, with correlations found between brain alignment and model size and world knowledge performance.
Yingzi Ma, Yulong Cao, Jiachen Sun +2 authors
Dolphins, a vision-language model enhanced with Grounded Chain of Thought, provides human-like capabilities as a conversational driving assistant by processing multimodal inputs and tailoring to specific driving tasks.
Ethan Weber, Aleksander Hołyński, Varun Jampani +4 authors
NeRFiller utilizes a 2D inpainting diffusion model to generate consistent 3D scene completions from sparse observations.
Shoufa Chen, Mengmeng Xu, Jiawei Ren +7 authors
GenTron, a Transformer-based diffusion model family, achieves superior visual quality and text alignment in image and video generation compared to existing methods.
Ryan Po, Guandao Yang, Kfir Aberman +1 authors
Orthogonal Adaptation enables efficient and scalable customization of text-to-image models by merging independently fine-tuned models without additional computational cost or loss of fidelity.
北京市昌平区探索星信息技术及软件开发工作室
京ICP备2026059466号