---
id: 20260804-T0-01
title: "新方法解决GPU分离推理的KV缓存传输难题"
title_en: "Topology-Aware Data Movement Solves Disaggregated GPU KV Cache Transfer"
url: https://ai.daily.yangsir.net/daily/20260804-T0-01
issue_date: 2026-08-04
publish_date: 2026-08-03T04:00:00.000Z
category: research
source_name: "arXiv cs.LG (ML)"
source_url: https://arxiv.org/abs/2607.28633
---

# 新方法解决GPU分离推理的KV缓存传输难题

一项新研究针对分离式 LLM 推理中的 KV 缓存传输问题提出拓扑感知数据移动方法。以 70B 模型为例，每个请求需传输 2.6GB 的 KV 缓存，现有系统无法有效处理。该方法根据网络拓扑优化数据路径，显著降低传输延迟，为大规模推理部署提供新方案。

## English Version

**Topology-Aware Data Movement Solves Disaggregated GPU KV Cache Transfer**

New research tackles KV cache transfer in disaggregated LLM inference, where a 70B model requires 2.6GB per request. The proposed topology-aware data movement method optimizes network paths, reducing transfer latency and offering a practical solution for large-scale inference deployments.

---

**来源**：[arXiv cs.LG (ML)](https://arxiv.org/abs/2607.28633)

**详情页**：https://ai.daily.yangsir.net/daily/20260804-T0-01

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*