---
id: 20260703-T0-01
title: "GRPO、Dr. GRPO和DAPO本质是同一个数字：组标准差恒等式"
title_en: "GRPO, Dr. GRPO, and DAPO Are All the Same Operation: Group-Standard-Deviation Identity"
url: https://ai.daily.yangsir.net/daily/20260703-T0-01
issue_date: 2026-07-03
publish_date: 2026-07-02T04:00:00.000Z
category: research
source_name: "arXiv cs.LG (ML)"
source_url: https://arxiv.org/abs/2607.00152
---

# GRPO、Dr. GRPO和DAPO本质是同一个数字：组标准差恒等式

最新arXiv研究揭示，GRPO、Dr. GRPO和DAPO这三种流行的大模型推理训练方法，表面上看似不同，实则都围绕同一个核心机制：组标准差（standard deviation），即模型对同一提示生成的不同回答之间的分歧程度。论文提出了统一的数学恒等式。

## English Version

**GRPO, Dr. GRPO, and DAPO Are All the Same Operation: Group-Standard-Deviation Identity**

A new arXiv paper reveals that GRPO, Dr. GRPO, and DAPO—three popular methods for training LLMs to reason—are mathematically equivalent. They all reduce to adjusting group standard deviation, measuring disagreement among sampled answers for a prompt.

---

**来源**：[arXiv cs.LG (ML)](https://arxiv.org/abs/2607.00152)

**详情页**：https://ai.daily.yangsir.net/daily/20260703-T0-01

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*