Sequential Functional Structured Tucker Compression for Large Language Model Attentions
A research paper proposes a new method for compressing large language model attentions, called FTC, which improves perplexity on several downstream tasks without requiring fine-tuning or gradient-based recovery.
Save an API key to vote.