TensorX
返回文献探索

Paper · arXiv 2411.09661

Adaptive Decoding via Latent Preference Optimization

Shehzaad Dhuliawala, Ilia Kulikov, Ping Yu, Asli Celikyilmaz, Jason Weston, Sainbayar Sukhbaatar, Jack Lanchantin

10 upvotesNovember 14, 2024arXiv 预印本
AI 摘要

Adaptive Decoding with Latent Preference Optimization dynamically adjusts sampling temperature during model inference, improving performance across tasks requiring varied temperature settings.

Adaptive DecodingLatent Preference Optimizationsampling temperatureinference timediscrete latent variablesUltraFeedbackCreative Story WritingGSM8K

Abstract

During language model decoding, it is known that using higher temperature sampling gives more creative responses, while lower temperatures are more factually accurate. However, such models are commonly applied to general instruction following, which involves both creative and fact seeking tasks, using a single fixed temperature across all examples and tokens. In this work, we introduce Adaptive Decoding, a layer added to the model to select the sampling temperature dynamically at inference time, at either the token or example level, in order to optimize performance. To learn its parameters we introduce Latent Preference Optimization (LPO) a general approach to train discrete latent variables such as choices of temperature. Our method outperforms all fixed decoding temperatures across a range of tasks that require different temperatures, including UltraFeedback, Creative Story Writing, and GSM8K.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Adaptive Decoding via Latent Preference Optimization | TensorX