Textbooks Are All You Need
Suriya Gunasekar, Yi Zhang, Jyoti Aneja +16 authors
A new compact Transformer-based large language model for code, phi-1, achieves high accuracy on coding benchmarks despite having fewer parameters than competing models.
Explore · 每周精选
发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。
Jade Copet, Felix Kreuk, Itai Gat +5 authors
MusicGen, a single-stage transformer language model, generates high-quality music conditioned on text or melodic features using efficient token interleaving patterns, outperforming existing models on text-to-music benchmarks.
50 篇论文 · 按点赞排序
Suriya Gunasekar, Yi Zhang, Jyoti Aneja +16 authors
A new compact Transformer-based large language model for code, phi-1, achieves high accuracy on coding benchmarks despite having fewer parameters than competing models.
Shuai Yang, Yifan Zhou, Ziwei Liu +1 authors
A novel framework adapts image diffusion models for video by generating key frames with hierarchical constraints and propagating them using patch matching and blending, achieving high-quality and temporally-coherent videos.
Ohad Rubin, Jonathan Berant
The Retrieval-Pretrained Transformer jointly trains a retrieval-augmented language model to model long texts by fusing retrieved information for improved prediction.
Luyang Zhu, Dawei Yang, Tyler Zhu +5 authors
A diffusion-based architecture unifies garment detail preservation and warping for pose and shape variation in virtual try-on tasks.
Paul K. Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen +27 authors
A unified multimodal language model combining text and speech capabilities outperforms existing systems in speech translation and zero-shot speech-to-text translation.
Shouyuan Chen, Sherman Wong, Liangjian Chen +1 authors
Position Interpolation extends the context window of RoPE-based LLMs up to 32768 tokens with minimal fine-tuning, maintaining task performance and stability compared to extrapolation methods.
Subhabrata Mukherjee, Arindam Mitra, Ganesh Jawahar +3 authors
Orca, a 13-billion parameter model, enhances small models by imitating the reasoning process from large foundation models using rich signals and diverse data, surpassing existing models in complex reasoning benchmarks.
Hugo Laurençon, Lucile Saulnier, Léo Tronchon +9 authors
The OBELICS dataset, containing interleaved image-text documents, is introduced and used to train competitive multimodal models.
Guilherme Penedo, Quentin Malartic, Daniel Hesslow +6 authors
Filtered and deduplicated web data alone can produce powerful large language models that outperform those trained on The Pile.
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng +10 authors
Using strong large language models as judges for evaluating other LLM-based chat assistants achieves high agreement with human preferences, offering a scalable and explainable solution compared to traditional benchmarks.
Minghua Liu, Chao Xu, Haian Jin +4 authors
A novel feed-forward method generates high-quality 360-degree 3D meshes from single images using a view-conditioned 2D diffusion model and an SDF-based neural surface reconstruction module.
Rishabh Agarwal, Nino Vieillard, Yongchao Zhou +4 authors
Generalized Knowledge Distillation addresses distribution mismatch in auto-regressive sequence models by training the student on its own output sequences and integrates with RL fine-tuning for effective distillation in tasks like summarization and instruction-tuning.
Kai Zhang, Lingbo Mo, Wenhu Chen +2 authors
MagicBrush, a large-scale manually annotated dataset for text-guided real image editing, enhances InstructPix2Pix model performance and highlights challenges in current baselines.
Zhiliang Peng, Wenhui Wang, Li Dong +4 authors
Kosmos-2, a Multimodal Large Language Model, combines text-to-visual grounding with existing MLLM capabilities, enabling tasks like referring expression comprehension and phrase grounding.
Xu Zhao, Wenchao Ding, Yongqi An +5 authors
A speed-up method using a CNN detector with an instance segmentation branch achieves comparable performance to SAM with 50 times faster runtime by training on a small subset of SAM's dataset.
Hadi Alzayer, Kevin Zhang, Brandon Feng +2 authors
A method for reconstructing 3D scenes beyond a camera's line of sight using eye reflections is proposed, refining cornea poses, radiance fields, and iris textures.
Ziyang Luo, Can Xu, Pu Zhao +7 authors
WizardCoder, a Code LLM fine-tuned with complex instructions using Evol-Instruct, outperforms other open-source and closed LLMs on several code generation benchmarks.
Yunpeng Bai, Xintao Wang, Yanpei Cao +3 authors
DreamDiffusion generates high-quality images from EEG signals using pre-trained text-to-image models and CLIP image encoder for improved alignment and robust representations, overcoming noise and data limitations.
Ali Hatamizadeh, Greg Heinrich, Hongxu Yin +4 authors
FasterViT, a hybrid CNN-ViT model, enhances CV tasks with high image throughput through Hierarchical Attention (HAT) and achieves state-of-the-art performance in accuracy versus throughput.
Kai Lv, Yuqing Yang, Tengxiao Liu +3 authors
A new optimizer reduces memory usage for full parameter fine-tuning of large language models, enabling 65B model training on a single machine with limited GPUs.
William Berrios, Gautam Mittal, Tristan Thrush +2 authors
LENS uses language models to reason over outputs from vision modules, achieving competitive performance in vision and vision-language tasks without multimodal training.
Biao Jiang, Xin Chen, Wen Liu +3 authors
MotionGPT, a unified motion-language model using discrete vector quantization, achieves top performance in various motion-related tasks by treating motion as a language similar to text.
Lionel Wong, Gabriel Grand, Alexander K. Lew +4 authors
A framework combining large language models and probabilistic programs translates language into a symbolic language of thought for rational, commonsense reasoning across multiple cognitive domains.
Difei Gao, Lei Ji, Luowei Zhou +4 authors
A multi-modal AI assistant, AssistGPT, with a Plan, Execute, Inspect, and Learn (PEIL) approach integrates LLMs with various tools to handle complex visual-based tasks and achieves state-of-the-art results on benchmarks and beyond.
Arnav Chavan, Zhuang Liu, Deepak Gupta +2 authors
GLoRA, an advanced method for parameter-efficient fine-tuning, enhances LoRA with a generalized prompt module and modular layer-wise structure search, offering superior performance across diverse tasks with fewer parameters and computational costs.
Zeju Qiu, Weiyang Liu, Haiwen Feng +6 authors
Orthogonal Finetuning and Constrained Orthogonal Finetuning methods enhance text-to-image diffusion models by preserving hyperspherical energy and improving stability, leading to better generation quality and speed.
Yuxian Gu, Li Dong, Furu Wei +1 authors
MiniLLM distills knowledge from large generative language models to smaller models using reverse KLD for better precision, quality, and performance.
Yuheng Li, Haotian Liu, Yangming Wen +1 authors
A text-to-image diffusion model is enhanced with adapter layers and regionally-guided sampling to achieve controlled generation of personalized objects with high fidelity.
Haocheng Xi, Changhao Li, Jianfei Chen +1 authors
A novel method for training transformers with 4-bit quantization achieves competitive accuracy and accelerates training on current GPUs.
北京市昌平区探索星信息技术及软件开发工作室
京ICP备2026059466号