---
id: 20260701-T0-13
title: "深度交错稀疏注意力：静态调度优于学习膨胀"
title_en: "Depth-Staggered Sparse Attention: Static Schedules Outperform Learned Methods"
url: https://ai.daily.yangsir.net/daily/20260701-T0-13
issue_date: 2026-07-01
publish_date: 2026-06-30T04:00:00.000Z
category: research
source_name: "arXiv cs.CL (NLP)"
source_url: https://arxiv.org/abs/2606.28560
---

# 深度交错稀疏注意力：静态调度优于学习膨胀

一项关于稀疏注意力的新研究提出了一种基于斐波那契间距的静态调度方案。该方案通过分层标量Alpha控制间距压缩，在21个语言模型的训练中表现优异。研究结论表明，这种静态稀疏注意力方案不仅能击败学习的膨胀方法，还能在密集注意力失效的场景下实现外推，有效提升了长上下文处理能力。

## English Version

**Depth-Staggered Sparse Attention: Static Schedules Outperform Learned Methods**

A new study proposes a sparse self-attention mechanism using Fibonacci-spaced offsets with depth-staggered static schedules. Results across 21 language models show that this static approach beats learned dilation and extrapolates effectively where dense attention fails.

---

**来源**：[arXiv cs.CL (NLP)](https://arxiv.org/abs/2606.28560)

**详情页**：https://ai.daily.yangsir.net/daily/20260701-T0-13

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*