NanoFlow: Towards Optimal Large Language Model Serving Throughput
Kan Zhu, Yilong Zhao, Liangyu Zhao +12 authors
NanoFlow enhances serving throughput by exploiting intra-device parallelism through operation co-scheduling and execution unit scheduling on GPUs.