ThinQuant: Scalable Rotation Learning for Weight and Activation Quantization of LLMs

ThinQuant is a new method for efficient rotation learning in large language models (LLMs). It reduces the computational cost of rotation learning by introducing a data selection procedure and an exact reduction of the optimization problem. This allows ThinQuant to scale to large architectures and achieve comparable performance to state-of-the-art methods like DartQuant and GPTAQ+QuaRoT.

RSS Score 0 9/30/2026, 4:00:00 AM Original Source
Save an API key to vote.