---
id: 20260923-T0-20
title: "TreeSpark：面向半自回归推测解码的校准负载自适应草稿树"
title_en: "TreeSpark: Calibrated Load-Adaptive Draft Trees for Semi-Autoregressive Speculative Decoding"
url: https://ai.daily.yangsir.net/daily/20260923-T0-20
issue_date: 2026-09-23
publish_date: 2026-09-22T04:00:00.000Z
category: research
source_name: "arXiv cs.CL (NLP)"
source_url: https://arxiv.org/abs/2609.22098
---

# TreeSpark：面向半自回归推测解码的校准负载自适应草稿树

arXiv新论文提出TreeSpark，针对推测解码中草稿模型生成候选token的效率问题。现有块草稿器通过单次骨干网络前向传播生成整块草稿token，TreeSpark进一步引入校准和负载自适应的草稿树结构，以提升半自回归推测解码的验证效率和推理加速效果。该方法旨在降低语言模型推理延迟，对需要高吞吐推理的LLM部署场景有潜在价值。

## English Version

**TreeSpark: Calibrated Load-Adaptive Draft Trees for Semi-Autoregressive Speculative Decoding**

A new arXiv paper introduces TreeSpark for semi-autoregressive speculative decoding. Existing block drafters make drafting nearly free by emitting an entire block of draft tokens in a single backbone pass. TreeSpark adds calibrated, load-adaptive draft trees to improve verification efficiency and inference speedup. The method aims to reduce language model inference latency, with potential value for high-throughput LLM deployment scenarios.

---

**来源**：[arXiv cs.CL (NLP)](https://arxiv.org/abs/2609.22098)

**详情页**：https://ai.daily.yangsir.net/daily/20260923-T0-20

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*