Efficient Long-context Language Model Training by Core Attention Disaggregation
Yonghao Zhuang, Junda Chen, Bo Pang +6 authors
CAD, a technique for long-context large language model training, improves throughput and balance by decoupling and distributing core attention computations.