TensorX
返回文献探索

Paper · arXiv 2310.15987

Dissecting In-Context Learning of Translations in GPTs

Vikas Raunak, Hany Hassan Awadalla, Arul Menezes

6 upvotesOctober 24, 2023arXiv 预印本
AI 摘要

Asymmetric perturbation of high-quality demonstrations in few-shot Machine Translation with LLMs shows that target-side perturbation significantly impacts translation quality, leading to the introduction of Zero-Shot-Context for enhanced zero-shot performance.

Large Language ModelsLLMsGPT-3Machine TranslationMTfew-shot samplesin-context learninghigh-quality demonstrationssource-target mappingstarget perturbationoutput text distributionZero-Shot-Contextzero-shot translationfew-shot prompted translations

Abstract

Most of the recent work in leveraging Large Language Models (LLMs) such as GPT-3 for Machine Translation (MT) has focused on selecting the few-shot samples for prompting. In this work, we try to better understand the role of demonstration attributes for the in-context learning of translations through perturbations of high-quality, in-domain demonstrations. We find that asymmetric perturbation of the source-target mappings yield vastly different results. We show that the perturbation of the source side has surprisingly little impact, while target perturbation can drastically reduce translation quality, suggesting that it is the output text distribution that provides the most important learning signal during in-context learning of translations. We propose a method named Zero-Shot-Context to add this signal automatically in Zero-Shot prompting. We demonstrate that it improves upon the zero-shot translation performance of GPT-3, even making it competitive with few-shot prompted translations.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Dissecting In-Context Learning of Translations in GPTs | TensorX