TensorX
返回文献探索

Paper · arXiv 2309.14592

Efficient Post-training Quantization with FP8 Formats

Haihao Shen, Naveen Mellempudi, Xin He, Qun Gao, Chang Wang, Mengni Wang

11 upvotesSeptember 26, 2023arXiv 预印本
AI 摘要

Research demonstrates that FP8 data formats improve quantization in deep learning models, outperforming INT8 in accuracy and suitability across various tasks.

LLMsDiffusion modelsFP8 data formatspost-training quantizationE5M2E4M3E3M4machine translationlanguage modelingtext generationimage classificationgenerationsegmentationworkload coverageIntel Neural Compressor

Abstract

Recent advances in deep learning methods such as LLMs and Diffusion models have created a need for improved quantization methods that can meet the computational demands of these modern architectures while maintaining accuracy. Towards this goal, we study the advantages of FP8 data formats for post-training quantization across 75 unique network architectures covering a wide range of tasks, including machine translation, language modeling, text generation, image classification, generation, and segmentation. We examine three different FP8 representations (E5M2, E4M3, and E3M4) to study the effects of varying degrees of trade-off between dynamic range and precision on model accuracy. Based on our extensive study, we developed a quantization workflow that generalizes across different network architectures. Our empirical results show that FP8 formats outperform INT8 in multiple aspects, including workload coverage (92.64% vs. 65.87%), model accuracy and suitability for a broader range of operations. Furthermore, our findings suggest that E4M3 is better suited for NLP models, whereas E3M4 performs marginally better than E4M3 on computer vision tasks. The code is publicly available on Intel Neural Compressor: https://github.com/intel/neural-compressor.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Efficient Post-training Quantization with FP8 Formats | TensorX