TensorX
返回文献探索

Paper · arXiv 2310.09753

When can transformers reason with abstract symbols?

Enric Boix-Adsera, Omid Saremi, Emmanuel Abbe, Samy Bengio, Etai Littwin, Joshua Susskind

3 upvotesOctober 15, 2023arXiv 预印本
AI 摘要

Transformers require vast amounts of data for regression tasks and exhibit decreased generalization with higher embedding dimensions in symbolic next-token-prediction tasks, but simple parameter modifications can improve their performance.

transformer large language modelsrelational reasoningnext-token-predictionembedding dimensiontrainable parameters

Abstract

We investigate the capabilities of transformer large language models (LLMs) on relational reasoning tasks involving abstract symbols. Such tasks have long been studied in the neuroscience literature as fundamental building blocks for more complex abilities in programming, mathematics, and verbal reasoning. For (i) regression tasks, we prove that transformers generalize when trained, but require astonishingly large quantities of training data. For (ii) next-token-prediction tasks with symbolic labels, we show an "inverse scaling law": transformers fail to generalize as their embedding dimension increases. For both settings (i) and (ii), we propose subtle transformer modifications which can reduce the amount of data needed by adding two trainable parameters per head.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
When can transformers reason with abstract symbols? | TensorX