ShamAN-Q: Shampoo Augmented NanoQuant for Sub-1-bit LLM Weights

ShamAN-Q is a sub-1-bit post-training quantization method for large language models (LLMs), building on the NanoQuant method. It uses a tractable dense curvature metric and Kullback-Leibler minimization to fit a Kronecker product to the empirical Fisher information matrix. ShamAN-Q improves perplexity on the WikiText-2 dataset and matches zero-shot accuracy on the Eleuther LM Evaluation Harness.

RSS Score 0 10/1/2026, 4:00:00 AM Original Source
Save an API key to vote.