Exponentially Faster Language Modelling
Peter Belcak, Roger Wattenhofer
FastBERT achieves significant inference speedup over BERT by selectively engaging a small fraction of neurons using fast feedforward networks.
Explore · 每周精选
发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。
Grégoire Mialon, Clémentine Fourrier, Craig Swift +3 authors
GAIA benchmarks general AI assistants using real-world questions that challenge both reasoning and multi-modality handling, showcasing a significant gap between human and AI performance.
44 篇论文 · 按点赞排序
Peter Belcak, Roger Wattenhofer
FastBERT achieves significant inference speedup over BERT by selectively engaging a small fraction of neurons using fast feedforward networks.
Arindam Mitra, Luciano Del Corro, Shweti Mahajan +12 authors
Orca 2 enhances smaller language models' reasoning abilities by teaching them diverse solution strategies, outperforming larger models on complex reasoning tasks.
Yan Zeng, Guoqiang Wei, Jiani Zheng +4 authors
PixelDance, a diffusion model-based approach, generates high-dynamic videos by incorporating image instructions for first and last frames alongside text instructions, surpassing current text-to-video methods in complexity and motion.
Vladimir Arkhipkin, Zein Shaheen, Viacheslav Vasilev +3 authors
A new two-stage latent diffusion model generates videos from text, achieving high quality and efficiency in keyframe synthesis and interpolation.
Jaeyoung Chung, Suyoung Lee, Hyeongjin Nam +2 authors
LucidDreamer generates domain-free 3D scenes using diffusion-based generative models and Gaussian splats, producing highly detailed results.
Bram Wallace, Meihua Dang, Rafael Rafailov +7 authors
A method called Diffusion-DPO aligns text-to-image diffusion models to human preferences using direct optimization on comparison data, improving visual appeal and prompt alignment.
Viraj Shah, Nataniel Ruiz, Forrester Cole +4 authors
ZipLoRA effectively combines independently trained style and subject LoRAs to enhance both subject and style fidelity in generative models.
Jason Weston, Sainbayar Sukhbaatar
System 2 Attention in Transformer-based Large Language Models improves factual accuracy and reduces bias by refining input context.
David Rein, Betty Li Hou, Asa Cooper Stickland +5 authors
A dataset of extremely difficult multiple-choice questions challenges both experts and AI systems, facilitating the development of scalable oversight methods for AI-generated knowledge.
Yiming Wang, Yu Lin, Xiaodong Zeng +1 authors
MultiLoRA improves multi-task adaptation for large language models by reducing the dominance of top singular vectors in LoRA parameter updates through horizontal scaling and modified initialization.
Di Chang, Yichun Shi, Quankai Gao +6 authors
MagicDance, a diffusion-based model, transfers human motion and facial expression across identities in dance videos, achieving high generalization and robust appearance control.
Antoine Guédon, Vincent Lepetit
A method for fast and precise mesh extraction from 3D Gaussian Splatting using a regularization term and Poisson reconstruction, enabling real-time editing and rendering quality.
Sang-Hoon Lee, Ha-Yeong Choi, Seung-Bin Kim +1 authors
HierSpeech++, an enhanced hierarchical variational autoencoder, achieves high-quality zero-shot speech synthesis by combining self-supervised representations and an efficient super-resolution framework.
Kai Yang, Jian Tao, Jiafei Lyu +6 authors
D3PO method directly fine-tunes diffusion models using human feedback without requiring a separate reward model, reducing computational cost and improving image quality.
Clifton Poth, Hannah Sterz, Indraneil Paul +7 authors
Adapters is an open-source library designed for parameter-efficient and modular transfer learning in large language models, offering a unified interface for various adapter methods.
Bin Lin, Bin Zhu, Yang Ye +3 authors
Video-LLaVA is a unified large vision-language model that enhances performance across various image and video benchmarks by integrating visual representations into the language feature space.
Animesh Sinha, Bo Sun, Anmol Kalia +14 authors
Style Tailoring is a finetuning method for Latent Diffusion Models that enhances the generation of sticker images, improving visual quality, prompt alignment, and scene diversity.
Shachar Rosenman, Vasudev Lal, Phillip Howard
An adaptive framework automatically enhances text-to-image prompts to improve generation quality using a pre-trained language model and constrained text decoding.
Vukasin Bozic, Danilo Dordervic, Daniele Coppola +1 authors
Shallow feed-forward networks can emulate the performance of the attention mechanism in Transformers, as demonstrated by their competitive results on sequence-to-sequence tasks using knowledge distillation.
Rohit Girdhar, Mannat Singh, Andrew Brown +7 authors
A text-to-video generation model that factors video creation into two steps—image generation followed by video generation—achieves superior quality and resolution without deep model cascades.
Rohit Gandikota, Joanna Materzynska, Tingrui Zhou +2 authors
Concept sliders enable precise control over image generation attributes using diffusion models by identifying low-rank parameter directions with minimal interference.
Tianyi Xie, Zeshun Zong, Yuxin Qiu +4 authors
PhysGaussian integrates Newtonian dynamics within 3D Gaussians for high-quality motion synthesis using a Material Point Method (MPM) that aligns with continuum mechanics principles, eliminating the need for traditional meshing techniques.
Yixun Liang, Xin Yang, Jiantao Lin +3 authors
Interval Score Matching improves text-to-3D generation by addressing Score Distillation Sampling's over-smoothing issue and using 3D Gaussian Splatting, achieving superior quality and efficiency.
Peng Wang, Hao Tan, Sai Bi +6 authors
A Pose-Free Large Reconstruction Model using self-attention for 3D reconstruction from few unposed images, estimating camera poses, and generalizing across datasets.
Hamish Ivison, Yizhong Wang, Valentina Pyatkin +8 authors
T\"ULU 2, an advanced suite of language models, achieves state-of-the-art performance through improved datasets, fine-tuning techniques, and direct preference optimization, surpassing even GPT-3.5-turbo-0301 on several benchmarks.
Feng Li, Qing Jiang, Hao Zhang +9 authors
A universal visual in-context prompting framework enhances zero-shot capabilities for both referring and generic vision tasks by using a versatile prompt encoder and reference image segments as context.
Shehan Munasinghe, Rusiru Thushara, Muhammad Maaz +4 authors
Video-LLaVA integrates audio transcriptions and uses a novel grounding module to enhance pixel-level grounding in large multimodal video models, leading to improved performance in generative and question-answering tasks.
Cicero Nogueira dos Santos, James Lee-Thorp, Isaac Noble +2 authors
A Mixture of Word Experts (MoWE) model demonstrates superior performance in NLP tasks compared to T5 models and traditional MoE models, using memory-augmented techniques and knowledge-rich routing functions.
Sai Saketh Rambhatla, Ishan Misra
SelfEval automates the assessment of text-image understanding in generative models by computing image likelihoods given text prompts, showing competitive performance to discriminative models on fine-grained tasks.
北京市昌平区探索星信息技术及软件开发工作室
京ICP备2026059466号