TensorX
返回文献探索

Paper · arXiv 2306.12509

Deep Language Networks: Joint Prompt Training of Stacked LLMs using Variational Inference

Alessandro Sordoni, Xingdi Yuan, Marc-Alexandre Côté, Matheus Pereira, Adam Trischler, Ziang Xiao, Arian Hosseini, Friederike Niedtner, Nicolas Le Roux

15 upvotesJune 21, 2023arXiv 预印本
AI 摘要

A Deep Language Network (DLN) achieves high performance by stacking language layers and optimizing prompts through variational inference, outperforming single layers and sometimes matching few-shot GPT-4.

Deep Language Network (DLN)language layerspromptsprompt optimizationvariational inferencefew-shot GPT-4

Abstract

We view large language models (LLMs) as stochastic language layers in a network, where the learnable parameters are the natural language prompts at each layer. We stack two such layers, feeding the output of one layer to the next. We call the stacked architecture a Deep Language Network (DLN). We first show how to effectively perform prompt optimization for a 1-Layer language network (DLN-1). We then show how to train 2-layer DLNs (DLN-2), where two prompts must be learnt. We consider the output of the first layer as a latent variable to marginalize, and devise a variational inference algorithm for joint prompt training. A DLN-2 reaches higher performance than a single layer, sometimes comparable to few-shot GPT-4 even when each LLM in the network is smaller and less powerful. The DLN code is open source: https://github.com/microsoft/deep-language-networks .

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号