TensorX
返回文献探索

Paper · arXiv 2410.20650

NeuZip: Memory-Efficient Training and Inference with Dynamic Compression of Neural Networks

Yongchang Hao, Yanshuai Cao, Lili Mou

17 upvotesOctober 28, 2024arXiv 预印本
AI 摘要

NeuZip compresses neural network weights using entropy for memory-efficient training and inference without performance loss.

NeuZipentropyweight compressionLlama-3memory footprinttraining dynamicsnear-lossless performance

Abstract

The performance of neural networks improves when more parameters are used. However, the model sizes are constrained by the available on-device memory during training and inference. Although applying techniques like quantization can alleviate the constraint, they suffer from performance degradation. In this work, we introduce NeuZip, a new weight compression scheme based on the entropy of floating-point numbers in neural networks. With NeuZip, we are able to achieve memory-efficient training and inference without sacrificing performance. Notably, we significantly reduce the memory footprint of training a Llama-3 8B model from 31GB to less than 16GB, while keeping the training dynamics fully unchanged. In inference, our method can reduce memory usage by more than half while maintaining near-lossless performance. Our code is publicly available.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号