---
id: 20260806-T0-07
title: "输出感知旋转：INT2 KV 缓存量化新方法，缓解长上下文推理瓶颈"
title_en: "Output-Aware Rotation: New INT2 KV-Cache Quantization Method"
url: https://ai.daily.yangsir.net/daily/20260806-T0-07
issue_date: 2026-08-06
publish_date: 2026-08-05T04:00:00.000Z
category: research
source_name: "arXiv cs.LG (ML)"
source_url: https://arxiv.org/abs/2608.02691
---

# 输出感知旋转：INT2 KV 缓存量化新方法，缓解长上下文推理瓶颈

该论文提出了输出感知旋转方法，用于改进 INT2 KV 缓存量化。KV 缓存是长上下文 LLM 推理中的主要内存和带宽瓶颈，现有的基于旋转的 INT2 方法虽优化了缓存统计特性，但在精度上仍有不足。新方法通过感知模型输出来调整旋转矩阵，在超低位宽下实现了更好的性能。

## English Version

**Output-Aware Rotation: New INT2 KV-Cache Quantization Method**

This paper proposes output-aware rotation for INT2 KV-cache quantization. KV cache is a major bottleneck in long-context LLM inference. Existing rotation-based INT2 methods optimize cache statistics but lack precision. The new method adjusts rotation matrices based on model output, achieving better performance at ultra-low bit widths.

---

**来源**：[arXiv cs.LG (ML)](https://arxiv.org/abs/2608.02691)

**详情页**：https://ai.daily.yangsir.net/daily/20260806-T0-07

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*