TryOnDiffusion: A Tale of Two UNets
Luyang Zhu, Dawei Yang, Tyler Zhu +5 authors
A diffusion-based architecture unifies garment detail preservation and warping for pose and shape variation in virtual try-on tasks.
Explore · 每周精选
发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。
Shuai Yang, Yifan Zhou, Ziwei Liu +1 authors
A novel framework adapts image diffusion models for video by generating key frames with hierarchical constraints and propagating them using patch matching and blending, achieving high-quality and temporally-coherent videos.
50 篇论文 · 按点赞排序
Luyang Zhu, Dawei Yang, Tyler Zhu +5 authors
A diffusion-based architecture unifies garment detail preservation and warping for pose and shape variation in virtual try-on tasks.
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng +10 authors
Using strong large language models as judges for evaluating other LLM-based chat assistants achieves high agreement with human preferences, offering a scalable and explainable solution compared to traditional benchmarks.
Hadi Alzayer, Kevin Zhang, Brandon Feng +2 authors
A method for reconstructing 3D scenes beyond a camera's line of sight using eye reflections is proposed, refining cornea poses, radiance fields, and iris textures.
Ziyang Luo, Can Xu, Pu Zhao +7 authors
WizardCoder, a Code LLM fine-tuned with complex instructions using Evol-Instruct, outperforms other open-source and closed LLMs on several code generation benchmarks.
Ali Hatamizadeh, Greg Heinrich, Hongxu Yin +4 authors
FasterViT, a hybrid CNN-ViT model, enhances CV tasks with high image throughput through Hierarchical Attention (HAT) and achieves state-of-the-art performance in accuracy versus throughput.
Difei Gao, Lei Ji, Luowei Zhou +4 authors
A multi-modal AI assistant, AssistGPT, with a Plan, Execute, Inspect, and Learn (PEIL) approach integrates LLMs with various tools to handle complex visual-based tasks and achieves state-of-the-art results on benchmarks and beyond.
Arnav Chavan, Zhuang Liu, Deepak Gupta +2 authors
GLoRA, an advanced method for parameter-efficient fine-tuning, enhances LoRA with a generalized prompt module and modular layer-wise structure search, offering superior performance across diverse tasks with fewer parameters and computational costs.
Zeju Qiu, Weiyang Liu, Haiwen Feng +6 authors
Orthogonal Finetuning and Constrained Orthogonal Finetuning methods enhance text-to-image diffusion models by preserving hyperspherical energy and improving stability, leading to better generation quality and speed.
Yuxian Gu, Li Dong, Furu Wei +1 authors
MiniLLM distills knowledge from large generative language models to smaller models using reverse KLD for better precision, quality, and performance.
George E. Dahl, Frank Schneider, Zachary Nado +22 authors
A new benchmark, AlgoPerf: Training Algorithms, addresses challenges in evaluating training algorithms by providing a competitive, time-to-result benchmark across workloads and optimizers.
Xiang Deng, Yu Gu, Boyuan Zheng +5 authors
Mind2Web, a dataset for generalist web agents, uses real-world websites and crowdsourced actions to improve LLM performance and generalizability.
Jifan Yu, Xiaozhi Wang, Shangqing Tu +32 authors
A benchmark for evaluating large language models emphasizes knowledge-related abilities, uses a mix of datasets, and incorporates self-contrast metrics to detect knowledge hallucination.
Arno Candel, Jon McKinney, Philipp Singer +12 authors
h2oGPT provides open-source, fine-tuned LLMs based on Generative Pretrained Transformers with 100% private document search capabilities.
Zhengyu Huang, Haoran Xie, Tsukasa Fukusato +1 authors
A latent space exploration of StyleGAN is used for generating high-quality anime portraits from incomplete sketches with a stroke-level disentanglement approach.
Weizhi Wang, Li Dong, Hao Cheng +4 authors
A framework called LongMem enables large language models to utilize long-term memory, overcoming input length limitations and improving performance on long-context tasks.
Dani Valevski, Danny Wasserman, Yossi Matias +1 authors
A method to condition text-to-image generation models on face embeddings in real-time, enhancing generative capabilities and bias mitigation.
Nikos Kolotouros, Thiemo Alldieck, Andrei Zanfir +3 authors
DreamHuman generates realistic animatable 3D human avatars using text by integrating text-to-image synthesis, neural radiance fields, and statistical human body models.
Tiange Luo, Chris Rockwell, Honglak Lee +1 authors
Cap3D generates high-quality descriptive text for 3D objects using pretrained models and datasets, surpassing human performance in quality, cost, and speed.
Chenyang Lyu, Minghao Wu, Longyue Wang +5 authors
Macaw-LLM integrates multimodal data including visual, audio, and text by employing a modality module, cognitive module, and novel alignment module, enhancing LLM capabilities across diverse data types.
Xiao Liu, Hanyu Lai, Hao Yu +6 authors
WebGLM is a web-enhanced question-answering system that improves on WebGPT by integrating web search, retrieval, and human preference to achieve better accuracy, efficiency, and cost-effectiveness.
Wenhao Yu, Nimrod Gileadi, Chuyuan Fu +17 authors
A new method uses large language models to define reward parameters for control policies, bridging high-level language instructions to low-level robotic actions through an interactive system.
Rithesh Kumar, Prem Seetharaman, Alejandro Luebs +2 authors
A universal neural audio compression algorithm achieves high fidelity at 8 kbps bandwidth by combining advancements in audio generation, vector quantization, and improved loss functions.
Xidong Feng, Yicheng Luo, Ziyan Wang +6 authors
ChessGPT integrates policy learning and language modeling by combining historical chess data and analytical insights to improve autonomous decision-making in chess games.
Michael Tschannen, Manoj Kumar, Andreas Steiner +3 authors
Captioning alone, when done with a carefully controlled comparison, proves to be as effective and sometimes more powerful than contrastive pretraining in vision and vision-language tasks.
Kush Bhatia, Avanika Narayan, Christopher De Sa +1 authors
TART improves large language models' reasoning capabilities using a task-agnostic Transformer-based module trained on synthetic logistic regression.
Jonathan Lorraine, Kevin Xie, Xiaohui Zeng +7 authors
A framework called Amortized Text-to-3D (ATT3D) efficiently generates 3D objects from text by sharing computation across multiple text prompts and enabling knowledge sharing for novel setups and smooth animations.
John J. Nay, David Karamardian, Sarah B. Lawsky +6 authors
LLMs demonstrate increasing legal understanding and accuracy in tax law when provided with additional context and prompting enhancements, though they have not yet reached the level of expert tax lawyers.
Jiatao Gu, Shuangfei Zhai, Yizhe Zhang +2 authors
A novel data-free distillation technique called BOOT accelerates diffusion models by learning a time-conditioned predictor, significantly improving generation speed without quality loss.
Laurynas Karazija, Iro Laina, Andrea Vedaldi +1 authors
Zero-shot open-vocabulary segmentation uses text-to-image diffusion models to sample support images, enhancing localization and background segmentation with pre-trained feature extractors.
北京市昌平区探索星信息技术及软件开发工作室
京ICP备2026059466号