Matryoshka Diffusion Models
Jiatao Gu, Shuangfei Zhai, Yizhe Zhang +2 authors
Matryoshka Diffusion Models use a NestedUNet architecture for joint denoising at multiple resolutions, enabling efficient high-resolution image and video synthesis.
Explore · 每周精选
发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。
Eyal Segalis, Dani Valevski, Danny Lumen +2 authors
Relabeling a text-to-image dataset with a specialized automatic captioning model improves image quality and semantic alignment in diffusion models.
50 篇论文 · 按点赞排序
Jiatao Gu, Shuangfei Zhai, Yizhe Zhang +2 authors
Matryoshka Diffusion Models use a NestedUNet architecture for joint denoising at multiple resolutions, enabling efficient high-resolution image and video synthesis.
Roee Hendel, Mor Geva, Amir Globerson
In-Context Learning in Large Language Models can be understood as compressing a training set into a task vector that modulates a transformer for output generation.
Aaron Gokaslan, A. Feder Cooper, Jasmine Collins +6 authors
A transfer learning technique and efficient training recipe using Creative Commons images enable high-quality text-to-image models with synthetic captions and reduced computational requirements.
Lianghui Zhu, Xinggang Wang, Xinlong Wang
Large Language Models fine-tuned as scalable judges (JudgeLM) achieve state-of-the-art performance in evaluating open-ended benchmarks through a comprehensive dataset and benchmark, enhancing judgment efficiency and accuracy.
Jingxiang Sun, Bo Zhang, Ruizhi Shao +4 authors
DreamCraft3D uses a hierarchical approach with diffusion models and tailored 3D priors to generate high-fidelity and coherent 3D objects from 2D references, addressing geometry consistency and texture fidelity.
Elias Frantar, Dan Alistarh
QMoE enables efficient execution of trillion-parameter MoE models on affordable hardware through compression and bespoke GPU decoding.
Fuxiao Liu, Tianrui Guan, Zongxia Li +4 authors
HallusionBench is a benchmark that highlights language hallucination and visual illusion issues in vision-language models (VLMs), showcasing the limitations of current state-of-the-art models like GPT-4V and LLaVA-1.5.
Joey Hejna, Rafael Rafailov, Harshit Sikchi +4 authors
A new regret-based algorithm, Contrastive Preference Learning (CPL), learns optimal policies directly from human preferences without learning a reward function, addressing optimization challenges in Reinforcement Learning from Human Feedback (RLHF).
Haoxiang Wang, Pavan Kumar Anasosalu Vasu, Fartash Faghri +6 authors
A unified model combining CLIP and SAM using multi-task learning, continual learning, and distillation techniques achieves state-of-the-art zero-shot semantic segmentation performance with reduced computational and data requirements.
Xiaoxiao Long, Yuan-Chen Guo, Cheng Lin +8 authors
Wonder3D leverages cross-domain diffusion models and geometry-aware normal fusion for high-quality, efficient, and consistent 3D mesh generation from single-view images.
Yang Wu, Shilong Wang, Hao Yang +4 authors
GPT-4V demonstrates strong visual understanding but has limitations in language comprehension, handling sensitive data, modalities like depth and audio, and fine visual nuances.
Samuel L. Smith, Andrew Brock, Leonard Berrada +1 authors
ConvNets pre-trained on a large dataset match the performance of Vision Transformers on ImageNet with comparable computational resources.
Ruida Wang, Wangchunshu Zhou, Mrinmaya Sachan
Synthesis Step by Step (S3) iteratively refines synthesized pseudo training examples using a large language model to reduce distributional discrepancies and improve small model performance on NLP tasks.
Shukang Yin, Chaoyou Fu, Sirui Zhao +7 authors
Woodpecker, a training-free method, corrects hallucinations in multimodal large language models by extracting key concepts, formulating questions, validating visual knowledge, generating visual claims, and correcting inconsistencies.
Kaiwen Zheng, Cheng Lu, Jianfei Chen +1 authors
A novel ODE solver for diffusion probabilistic models optimizes sampling efficiency and sample quality by minimizing discretization errors and introducing empirical model statistics, achieving faster and better results in both pixel-space and latent-space models.
Changli Tang, Wenyi Yu, Guangzhi Sun +6 authors
SALMONN, an integrated multimodal model combining a pre-trained text-based LLM with speech and audio encoders, demonstrates competitive performance and emergent abilities in various speech and audio tasks.
Zhaoyang Wang, Shaohan Huang, Yuxuan Liu +8 authors
A tailored learning approach uses interactive multi-round distillation from large language models to enhance reasoning abilities in smaller LMs, fostering democratization through self-reflection and customized training.
Sudarshan Babu, Richard Liu, Avery Zhou +3 authors
HyperFields, a method using a dynamic hypernetwork and NeRF distillation, generates text-conditioned NeRFs efficiently across multiple scenes with potentially fast fine-tuning.
Yuchen Zhuang, Xiang Chen, Tong Yu +5 authors
ToolChain*, an A* search-based planning algorithm, enhances the efficiency of LLM-based autonomous agents in navigating expansive action spaces by pruning incorrect API function calls.
Zichang Liu, Jue Wang, Tri Dao +8 authors
DejaVu system predicts and exploits contextual sparsity on large language models to significantly reduce inference latency without affecting quality.
Sidharth Mudgal, Jong Lee, Harish Ganapathy +10 authors
Controlled decoding is an off-policy reinforcement learning method that uses a prefix scorer to guide language model generation towards high rewards and can handle multiple objectives without additional complexity.
Shih-yang Liu, Zechun Liu, Xijie Huang +2 authors
LLM-FP4 quantizes both weights and activations of large language models to 4-bit floating-point, achieving near-full-precision performance with post-training quantization.
Kevin Lin, Zhengyuan Yang, Linjie Li +2 authors
DEsignBench evaluates text-to-image models in visual design contexts with human and automatic assessments across various design criteria.
Bangbang Yang, Wenqi Dong, Lin Ma +4 authors
A novel indoor scene texturing framework uses diffusion-based methods for text-driven texture generation with high detail and spatial coherence, addressing challenges in 3D spatial applications like XR/VR through dual texture alignment and a separated inpainting strategy.
Xiao Yu, Baolin Peng, Michel Galley +2 authors
TriPosT trains smaller models to self-improve by interacting with larger language models to collect feedback, narrowing the performance gap in math and reasoning tasks.
Zhihan Zhang, Shuohang Wang, Wenhao Yu +6 authors
Auto-Instruct automatically improves the quality of instructions for large language models by generating and ranking diverse candidates using a scoring model trained on various NLP tasks.
Weijia Shi, Anirudh Ajith, Mengzhou Xia +5 authors
Researchers propose Min-K% Prob, a method for detecting pretraining data in large language models without needing knowledge of the pretraining corpus, using a benchmark that supports gold truth detection.
Guanzheng Chen, Xin Li, Zaiqiao Meng +2 authors
CLEX, a method for continuous length extrapolation, extends the context window of Transformer-based LLMs beyond the training length with minimal performance loss.
Saurabh Garg, Mehrdad Farajtabar, Hadi Pouransari +5 authors
Web-scale Time-Continual (TiC) benchmarks are introduced to evaluate and improve the temporal robustness of vision-language models through efficient continual learning methods.
北京市昌平区探索星信息技术及软件开发工作室
京ICP备2026059466号