Large Language Models as Optimizers
Chengrun Yang, Xuezhi Wang, Yifeng Lu +4 authors
OPRO, a method using large language models to optimize tasks described in natural language, outperforms human-designed prompts on various benchmark datasets.
Explore · 每周精选
发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。
403 篇论文 · 按点赞排序
Chengrun Yang, Xuezhi Wang, Yifeng Lu +4 authors
OPRO, a method using large language models to optimize tasks described in natural language, outperforms human-designed prompts on various benchmark datasets.
Anton Razzhigaev, Arseniy Shakhmatov, Anastasia Maltseva +7 authors
Kandinsky1, a latent diffusion architecture, achieves high-quality text-to-image generation by integrating image prior models and modified MoVQ autoencoders, outperforming other open-source models.
Arindam Mitra, Luciano Del Corro, Shweti Mahajan +12 authors
Orca 2 enhances smaller language models' reasoning abilities by teaching them diverse solution strategies, outperforming larger models on complex reasoning tasks.
Akio Kodaira, Chenfeng Xu, Toshiki Hazama +7 authors
StreamDiffusion, a real-time diffusion pipeline, enhances interactive image generation by implementing batching denoising, residual classifier-free guidance, and stochastic similarity filtering, which boost throughput and reduce energy consumption.
Luyang Zhu, Dawei Yang, Tyler Zhu +5 authors
A diffusion-based architecture unifies garment detail preservation and warping for pose and shape variation in virtual try-on tasks.
Mukul Singh, José Cambronero, Sumit Gulwani +3 authors
CodeFusion, a diffusion model for code generation, outperforms auto-regressive models in natural language to code tasks by iteratively refining the entire program.
Yan Zeng, Guoqiang Wei, Jiani Zheng +4 authors
PixelDance, a diffusion model-based approach, generates high-dynamic videos by incorporating image instructions for first and last frames alongside text instructions, surpassing current text-to-video methods in complexity and motion.
Vincent Perot, Kai Kang, Florian Luisier +7 authors
LMDX adapts large language models for document information extraction with grounding guarantees, setting a new state-of-the-art on VRDU and CORD benchmarks.
Chenyang Si, Ziqi Huang, Yuming Jiang +1 authors
A method called FreeU improves diffusion U-Net models' generation quality by re-weighting skip connections and backbone features without additional training.
Yuwei Guo, Ceyuan Yang, Anyi Rao +4 authors
A framework inserts a motion model into existing text-to-image models to enable animation, using video clips to learn motion priors while preserving model domain and diversity.
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman +1 authors
QLoRA enables efficient finetuning of large language models using 4-bit quantization and Low Rank Adapters, achieving high performance with reduced memory usage.
Dahyun Kim, Chanjun Park, Sanghoon Kim +15 authors
A novel technique, depth up-scaling (DUS), efficiently enhances large language models (LLMs) without complex changes, building SOLAR 10.7B that outperforms existing open-source models in various NLP tasks, including instruction-following.
Junsong Chen, Jincheng Yu, Chongjian Ge +8 authors
PIXART-$\alpha$ is a Transformer-based text-to-image diffusion model that achieves high-quality image synthesis with low training costs using advanced training strategies and efficient architecture.
Zhen Li, Mingdeng Cao, Xintao Wang +3 authors
PhotoMaker is an efficient personalized text-to-image generation method that encodes ID images into a unified embedding, achieving high ID fidelity, speed improvements, and generalization capabilities with an ID-oriented data pipeline.
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang +6 authors
PagedAttention algorithm and vLLM system enhance the throughput of large language models by efficiently managing memory and reducing waste in the key-value cache.
Chunyi Sun, Junlin Han, Weijian Deng +3 authors
3D-GPT leverages large language models to drive 3D modeling tasks, enhancing scene descriptions and interfacing with 3D software through integrated agents.
Runpei Dong, Chunrui Han, Yuang Peng +11 authors
DreamLLM, a framework for Multimodal Large Language Models, directly samples in the multimodal space to enhance comprehension and creation synergy, enabling free-form interleaved content generation.
Omri Avrahami, Amir Hertz, Yael Vinker +5 authors
A new automated method generates consistent characters from text prompts by iteratively refining a set of similar images.
Michal Geyer, Omer Bar-Tal, Shai Bagon +1 authors
A framework for text-driven video editing uses diffusion models to ensure consistency in the edited video by propagating features based on inter-frame correspondences.
Vladimir Arkhipkin, Zein Shaheen, Viacheslav Vasilev +3 authors
A new two-stage latent diffusion model generates videos from text, achieving high quality and efficiency in keyframe synthesis and interpolation.
Zhengzhong Liu, Aurick Qiao, Willie Neiswanger +25 authors
LLM360 initiative promotes full transparency and reproducibility in LLM training by open-sourcing training code, data, model checkpoints, and intermediate results.
Seungone Kim, Jamin Shin, Yejin Cho +8 authors
Prometheus, an open-source LLM, matches GPT-4's evaluation capabilities using customized score rubrics and outperforms ChatGPT across various benchmarks.
Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster +6 authors
Llemma, a large language model pretrained on mathematical data, outperforms existing models and demonstrates tool use and formal theorem proving capabilities.
Seungwhan Moon, Andrea Madotto, Zhaojiang Lin +10 authors
AnyMAL is a unified model that processes multiple modalities (text, image, video, audio, IMU) and generates text, achieving top performance in multimodal tasks through a pre-trained aligner and fine-tuning with diverse instructions.
Paul K. Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen +27 authors
A unified multimodal language model combining text and speech capabilities outperforms existing systems in speech translation and zero-shot speech-to-text translation.
Sanchit Gandhi, Patrick von Platen, Alexander M. Rush
Distil-Whisper, a smaller and faster variant of the Whisper model, achieves nearly the same performance with fewer resources and is optimized for low-latency environments.
Tengchao Lv, Yupan Huang, Jingye Chen +11 authors
Kosmos-2.5, a unified multimodal model, generates spatially-aware and structured text from text-intensive images using a Transformer architecture and task-specific prompts.
Zhengqi Li, Richard Tucker, Noah Snavely +1 authors
A frequency-coordinated diffusion sampling process is used to predict long-term motion representations for still images, enabling dynamic video creation and interactive scene manipulation.
Xueyao Zhang, Liumeng Xue, Yuancheng Wang +10 authors
Amphion is a toolkit for audio, music, and speech generation that includes model visualizations, vocoders, and evaluation metrics to support reproducible research and training for researchers.
Jon Saad-Falcon, Joe Barrow, Alexa Siu +3 authors
PDFTriage enables LLMs to retrieve context effectively from structured documents like PDFs and web pages for better document QA performance.
北京市昌平区探索星信息技术及软件开发工作室
京ICP备2026059466号