ThinQuant: Scalable Rotation Learning for Weight and Activation Quantization of LLMs
ThinQuant is a new method for efficient rotation learning in large language models (LLMs). It reduces the computational cost of rotation learning by introducing a data selection procedure and an exact reduction of the optimization problem. This allows ThinQuant to scale to large architectures and achieve comparable performance to state-of-the-art methods like DartQuant and GPTAQ+QuaRoT.
Save an API key to vote.