TensorX
返回文献探索

Paper · arXiv 2406.00888

Show, Don't Tell: Aligning Language Models with Demonstrated Feedback

Omar Shaikh, Michelle Lam, Joey Hejna, Yijia Shao, Michael Bernstein, Diyi Yang

31 upvotesJune 2, 2024arXiv 预印本
AI 摘要

Demonstration ITerated Task Optimization (DITTO) aligns language models to specific settings using very few demonstrations, outperforming few-shot prompting and supervised fine-tuning across various domains.

LLMssupervised finetuningRLHFDITTOonline imitation learningfew-shot promptingtask alignmentnews articlesemailsblog postsuser studywin-ratesself-play methods

Abstract

Language models are aligned to emulate the collective voice of many, resulting in outputs that align with no one in particular. Steering LLMs away from generic output is possible through supervised finetuning or RLHF, but requires prohibitively large datasets for new ad-hoc tasks. We argue that it is instead possible to align an LLM to a specific setting by leveraging a very small number (<10) of demonstrations as feedback. Our method, Demonstration ITerated Task Optimization (DITTO), directly aligns language model outputs to a user's demonstrated behaviors. Derived using ideas from online imitation learning, DITTO cheaply generates online comparison data by treating users' demonstrations as preferred over output from the LLM and its intermediate checkpoints. We evaluate DITTO's ability to learn fine-grained style and task alignment across domains such as news articles, emails, and blog posts. Additionally, we conduct a user study soliciting a range of demonstrations from participants (N=16). Across our benchmarks and user study, we find that win-rates for DITTO outperform few-shot prompting, supervised fine-tuning, and other self-play methods by an average of 19% points. By using demonstrations as feedback directly, DITTO offers a novel method for effective customization of LLMs.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号