TensorX
返回文献探索

Paper · arXiv 2311.06243

Parameter-Efficient Orthogonal Finetuning via Butterfly Factorization

Weiyang Liu, Zeju Qiu, Yao Feng, Yuliang Xiu, Yuxuan Xue, Longhui Yu, Haiwen Feng, Zhen Liu, Juyeon Heo, Songyou Peng, Yandong Wen, Michael J. Black, Adrian Weller, Bernhard Schölkopf

18 upvotesNovember 10, 2023arXiv 预印本
AI 摘要

A novel parameter-efficient finetuning method, Orthogonal Butterfly (BOFT), is proposed to enhance the generalization and reduce the computational cost of adapting large foundation models to downstream tasks.

Orthogonal Finetuning (OFT)parameter-efficiencybutterfly structuresOrthogonal Butterfly (BOFT)generalized orthogonal finetuning frameworkvision transformerslanguage modelstext-to-image diffusion models

Abstract

Large foundation models are becoming ubiquitous, but training them from scratch is prohibitively expensive. Thus, efficiently adapting these powerful models to downstream tasks is increasingly important. In this paper, we study a principled finetuning paradigm -- Orthogonal Finetuning (OFT) -- for downstream task adaptation. Despite demonstrating good generalizability, OFT still uses a fairly large number of trainable parameters due to the high dimensionality of orthogonal matrices. To address this, we start by examining OFT from an information transmission perspective, and then identify a few key desiderata that enable better parameter-efficiency. Inspired by how the Cooley-Tukey fast Fourier transform algorithm enables efficient information transmission, we propose an efficient orthogonal parameterization using butterfly structures. We apply this parameterization to OFT, creating a novel parameter-efficient finetuning method, called Orthogonal Butterfly (BOFT). By subsuming OFT as a special case, BOFT introduces a generalized orthogonal finetuning framework. Finally, we conduct an extensive empirical study of adapting large vision transformers, large language models, and text-to-image diffusion models to various downstream tasks in vision and language.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Parameter-Efficient Orthogonal Finetuning via Butterfly Factorization | TensorX