Hugging Face Blog·· 2025-01-23AI 评分38
NVIDIA KVPress 如何压缩 KV Cache 实现长上下文 LLM
Mastering Long Contexts in LLMs with KVPress
AI 导读
NVIDIA 推出 Python 工具包 KVPress,通过多种压缩算法降低长上下文 LLM 的 KV Cache 内存占用。以 Llama 3-70B 在 bfloat16 下处理 1M token 为例,KV Cache 需 327.6GB,占约 470GB 总内存的 70%。
来源:Hugging Face Blog · huggingface.co