Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length
Xuezhe Ma, Xiaomeng Yang, Wenhan Xiong +7 authors
Megaldon, an enhanced neural architecture, provides efficient long-sequence modeling by combining and improving upon components from mega and transformer models, achieving better pretraining efficiency and performance.