LongNet: Scaling Transformers to 1,000,000,000 Tokens
Jiayu Ding, Shuming Ma, Li Dong +4 authors
LongNet, a Transformer variant with dilated attention, enables scaling to over 1 billion tokens with linear computational complexity and maintains performance on shorter sequences.