TensorX
返回文献探索

Paper · arXiv 2607.07964

KronQ: LLM Quantization via Kronecker-Factored Hessian

Donghyun Lee, Yuhang Li, Ruokai Yin, Priyadarshini Panda

33 upvotesJuly 8, 2026arXiv 预印本
AI 摘要

KronQ improves post-training quantization by incorporating gradient covariance via Kronecker-factored Hessian approximations, enabling bidirectional incoherence processing and gradient-aware mixed-precision allocation.

post-training quantizationsecond-order PTQGPTQKronecker-factored Hessiangradient covariancebidirectional incoherence processingrandom rotationmixed-precision allocationweight-only quantization

Abstract

Post-training quantization (PTQ) is a widely adopted technique for compressing large language models (LLMs) without retraining. Existing second-order PTQ methods, including GPTQ, construct quantization objectives exclusively from input activation statistics, effectively assuming that all output channels contribute equally to the layer-wise reconstruction objective. We propose KronQ, a PTQ framework that challenges this assumption by introducing the gradient covariance into the quantization pipeline. Under the Kronecker-factored Hessian approximation, the quantization loss depends jointly on both the activation and gradient covariances, and KronQ exploits this at two complementary levels. (1) KronQ introduces bidirectional incoherence processing, extending the existing input-side random rotation to the output dimension using the gradient covariance, reducing weight magnitude variance across both input and output dimensions. (2) KronQ derives a new sensitivity metric for inter-layer mixed-precision allocation, driven by the gradient and activation Hessian traces. Notably, in the case of 2-bit weight-only quantization on LLaMA-3-70B, while GPTQ and GPTAQ diverge or produce degenerate quantizations (>2000 perplexity on WikiText-2), KronQ achieves 7.93 perplexity.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
KronQ: LLM Quantization via Kronecker-Factored Hessian | TensorX