TensorX
返回文献探索

Paper · arXiv 2306.15400

Length Generalization in Arithmetic Transformers

Samy Jelassi, Stéphane d'Ascoli, Carles Domingo-Enrich, Yuhuai Wu, Yuanzhi Li, François Charton

4 upvotesJune 27, 2023arXiv 预印本
AI 摘要

Transformers use relative position embeddings for length generalization in simple tasks like addition and propose train set priming to achieve generalization in complex tasks like multiplication.

transformersrelative position embeddingslength generalizationadditionmultiplicationtrain set priming

Abstract

We examine how transformers cope with two challenges: learning basic integer arithmetic, and generalizing to longer sequences than seen during training. We find that relative position embeddings enable length generalization for simple tasks, such as addition: models trained on 5-digit numbers can perform 15-digit sums. However, this method fails for multiplication, and we propose train set priming: adding a few (10 to 50) long sequences to the training set. We show that priming allows models trained on 5-digit times 3-digit multiplications to generalize to 35times 3 examples. We also show that models can be primed for different generalization lengths, and that the priming sample size scales as the logarithm of the training set size. Finally, we discuss potential applications of priming beyond arithmetic.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号