Exponentially Faster Language Modelling
Peter Belcak, Roger Wattenhofer
FastBERT achieves significant inference speedup over BERT by selectively engaging a small fraction of neurons using fast feedforward networks.
Explore · 每周精选
发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。
Grégoire Mialon, Clémentine Fourrier, Craig Swift +3 authors
GAIA benchmarks general AI assistants using real-world questions that challenge both reasoning and multi-modality handling, showcasing a significant gap between human and AI performance.
50 篇论文 · 按点赞排序
Peter Belcak, Roger Wattenhofer
FastBERT achieves significant inference speedup over BERT by selectively engaging a small fraction of neurons using fast feedforward networks.
Bin Xiao, Haiping Wu, Weijian Xu +6 authors
A new prompt-based vision foundation model, Florence-2, is introduced for diverse vision and vision-language tasks, achieving strong zero-shot and fine-tuning capabilities with comprehensive annotations.
Simian Luo, Yiqin Tan, Suraj Patil +6 authors
LCMs enhance text-to-image generation by distilling from LDMs using LoRA for reduced memory and superior quality, and introduce LCM-LoRA as a plug-in accelerator for various tasks.
Arindam Mitra, Luciano Del Corro, Shweti Mahajan +12 authors
Orca 2 enhances smaller language models' reasoning abilities by teaching them diverse solution strategies, outperforming larger models on complex reasoning tasks.
Yan Zeng, Guoqiang Wei, Jiani Zheng +4 authors
PixelDance, a diffusion model-based approach, generates high-dynamic videos by incorporating image instructions for first and last frames alongside text instructions, surpassing current text-to-video methods in complexity and motion.
Vladimir Arkhipkin, Zein Shaheen, Viacheslav Vasilev +3 authors
A new two-stage latent diffusion model generates videos from text, achieving high quality and efficiency in keyframe synthesis and interpolation.
Omri Avrahami, Amir Hertz, Yael Vinker +5 authors
A new automated method generates consistent characters from text prompts by iteratively refining a set of similar images.
Sanchit Gandhi, Patrick von Platen, Alexander M. Rush
Distil-Whisper, a smaller and faster variant of the Whisper model, achieves nearly the same performance with fewer resources and is optimized for low-latency environments.
Jaeyoung Chung, Suyoung Lee, Hyeongjin Nam +2 authors
LucidDreamer generates domain-free 3D scenes using diffusion-based generative models and Gaussian splats, producing highly detailed results.
Shilong Liu, Hao Cheng, Haotian Liu +10 authors
LLaVA-Plus, a general-purpose multimodal assistant, enhances large multimodal models by integrating pre-trained vision and vision-language models, performing tool-assisted tasks and improving interaction through direct image grounding.
Yicong Hong, Kai Zhang, Jiuxiang Gu +7 authors
A Large Reconstruction Model using a transformer-based architecture predicts 3D neural radiance fields from single images using massive multi-view training data.
Bram Wallace, Meihua Dang, Rafael Rafailov +7 authors
A method called Diffusion-DPO aligns text-to-image diffusion models to human preferences using direct optimization on comparison data, improving visual appeal and prompt alignment.
Viraj Shah, Nataniel Ruiz, Forrester Cole +4 authors
ZipLoRA effectively combines independently trained style and subject LoRAs to enhance both subject and style fidelity in generative models.
Yanwu Xu, Yang Zhao, Zhisheng Xiao +1 authors
UFOGen, a hybrid diffusion-GAN model, achieves efficient one-step text-to-image synthesis at high quality.
Wojciech Zielonka, Timur Bagautdinov, Shunsuke Saito +3 authors
A new 3D controllable avatar model uses Gaussian splats for photorealistic rendering in real-time, employing cage deformations driven by joint angles and keypoints, outperforming existing methods.
Ming Li, Pan Zhou, Jia-Wei Liu +4 authors
A framework named Instant3D generates 3D objects from text prompts in under one second using a novel network and adaptive algorithms to enhance efficiency and quality.
Shih-Lun Wu, Chris Donahue, Shinji Watanabe +1 authors
Music ControlNet, a diffusion-based model, introduces precise, time-varying control over audio generation, offering enhanced realism and efficiency compared to existing models.
Jason Weston, Sainbayar Sukhbaatar
System 2 Attention in Transformer-based Large Language Models improves factual accuracy and reduces bias by refining input context.
Wei-Ge Chen, Irina Spiridonova, Jianwei Yang +2 authors
LLaVA-Interactive is a cost-efficient multimodal dialogue system that integrates visual chat, image segmentation, and generation capabilities to enhance human-AI interaction.
Minghua Liu, Ruoxi Shi, Linghao Chen +7 authors
One-2-3-45++ generates detailed 3D meshes from a single image using fine-tuned 2D diffusion models and multi-view conditioned 3D diffusion models, balancing speed and quality.
David Rein, Betty Li Hou, Asa Cooper Stickland +5 authors
A dataset of extremely difficult multiple-choice questions challenges both experts and AI systems, facilitating the development of scalable oversight methods for AI-generated knowledge.
Ke Hong, Guohao Dai, Jiaming Xu +6 authors
FlashDecoding++ addresses challenges in LLM inference through asynchronized softmax, flat GEMM optimization, and heuristic dataflow adaptation, achieving significant speedups.
Yiming Wang, Yu Lin, Xiaodong Zeng +1 authors
MultiLoRA improves multi-task adaptation for large language models by reducing the dominance of top singular vectors in LoRA parameter updates through horizontal scaling and modified initialization.
Zihao Wang, Shaofei Cai, Anji Liu +9 authors
JARVIS-1, an open-world agent in Minecraft, uses multimodal perception and a memory system to perform complex tasks and improve over time.
Hanoona Rasheed, Muhammad Maaz, Sahal Shaji +7 authors
GLaMM is a multimodal model that generates visually grounded language responses with object segmentation masks from both text and optional visual prompts.
Meredith Ringel Morris, Jascha Sohl-dickstein, Noah Fiedel +5 authors
The framework proposes a classification system for AGI capabilities and behavior using levels of performance and generality, with principles to assess risks and measure progress.
Yew Ken Chia, Guizhen Chen, Luu Anh Tuan +2 authors
Contrastive chain of thought, utilizing both valid and invalid reasoning examples, improves language model reasoning and generalization compared to conventional methods.
Yilin Zhao, Xinbin Yuan, Shanghua Gao +4 authors
A framework for generating anthropomorphized personas with diverse voices and appearances from text descriptions using LLMs and generative models, with improved face landmark detection for automatic animation.
Bo Li, Peiyuan Zhang, Jingkang Yang +3 authors
OtterHD-8B, an advanced multimodal model from Fuyu-8B, excels in processing high-resolution inputs and discerning detailed spatial relationships through MagnifierBench, an evaluation framework highlighting the importance of vision encoder flexibility.
北京市昌平区探索星信息技术及软件开发工作室
京ICP备2026059466号