Mixture-of-Depths: Dynamically allocating compute in transformer-based language models
David Raposo, Sam Ritter, Blake Richards +3 authors
Transformer-based models can dynamically allocate computational resources to specific positions in input sequences, leading to efficient performance with reduced FLOPs.