TensorX
返回文献探索

Paper · arXiv 2311.11829

System 2 Attention (is something you might need too)

Jason Weston, Sainbayar Sukhbaatar

43 upvotesNovember 20, 2023arXiv 预印本
AI 摘要

System 2 Attention in Transformer-based Large Language Models improves factual accuracy and reduces bias by refining input context.

soft attentionTransformer-based Large Language Models (LLMs)System 2 Attention (S2A)natural language reasoningcontext regenerationfactualityobjectivitysycophancyQAmath word problemslongform generation

Abstract

Soft attention in Transformer-based Large Language Models (LLMs) is susceptible to incorporating irrelevant information from the context into its latent representations, which adversely affects next token generations. To help rectify these issues, we introduce System 2 Attention (S2A), which leverages the ability of LLMs to reason in natural language and follow instructions in order to decide what to attend to. S2A regenerates the input context to only include the relevant portions, before attending to the regenerated context to elicit the final response. In experiments, S2A outperforms standard attention-based LLMs on three tasks containing opinion or irrelevant information, QA, math word problems and longform generation, where S2A increases factuality and objectivity, and decreases sycophancy.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号