KV-Kaizen: Learning Context-Adaptive Cache Compression Choices
The paper proposes KV-Kaizen, a method for learning context-adaptive cache compression choices for Large Language Models (LLMs). This approach can reduce memory usage without compromising accuracy, enabling the use of larger models and improving inference performance.
Save an API key to vote.