---
id: 20260905-T0-04
title: "GrowPage：按需分配KV预算，高效支撑长输出推理"
title_en: "GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning"
url: https://ai.daily.yangsir.net/daily/20260905-T0-04
issue_date: 2026-09-05
publish_date: 2026-09-04T04:00:00.000Z
category: research
source_name: "arXiv cs.AI"
source_url: https://arxiv.org/abs/2609.03494
---

# GrowPage：按需分配KV预算，高效支撑长输出推理

一项新研究提出GrowPage方法，用于解决长输出推理场景下KV缓存带来的内存瓶颈问题。与现有需要预设每请求预算的压缩方法不同，GrowPage按需调整KV预算，从而在不牺牲性能的前提下提高LLM服务的内存效率和吞吐量。

## English Version

**GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning**

New research introduces GrowPage, a method addressing KV cache memory bottlenecks in long-output reasoning. Unlike existing compression techniques that rely on predefined per-request budgets, GrowPage adjusts KV budgets on-demand, improving memory efficiency and throughput in LLM serving without performance compromise.

---

**来源**：[arXiv cs.AI](https://arxiv.org/abs/2609.03494)

**详情页**：https://ai.daily.yangsir.net/daily/20260905-T0-04

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*