---
id: 20260729-T0-06
title: "PGPO新方法提升LLM Agent在长期任务中的成功率"
title_en: "PGPO Boosts LLM Agent Success Rates on Long-Horizon Tasks"
url: https://ai.daily.yangsir.net/daily/20260729-T0-06
issue_date: 2026-07-29
publish_date: 2026-07-28T04:00:00.000Z
category: research
source_name: "arXiv cs.LG (ML)"
source_url: https://arxiv.org/abs/2607.22724
---

# PGPO新方法提升LLM Agent在长期任务中的成功率

arXiv论文提出Progress-conditioned Group Policy Optimization（PGPO），一种改进的群体策略优化方法，用于训练LLM Agent处理需要多步推理的长期任务。传统对比方法在复杂任务中因“组内其他轨迹表现太差”而难以给出有效比较信号。PGPO通过引入进度条件（根据已完成的子任务比例评估）来优化对比，显著提升了Agent在困难任务中的成功率。对自动编程、多步工具调用等场景有实际价值。

## English Version

**PGPO Boosts LLM Agent Success Rates on Long-Horizon Tasks**

A new arXiv paper introduces Progress-conditioned Group Policy Optimization (PGPO), which improves LLM agent training for long-horizon tasks. Traditional group-based methods fail when comparison trajectories are too poor. PGPO adds a progress condition based on sub-task completion ratio, significantly boosting agent success rates in complex tasks like automated coding and multi-step tool use.

---

**来源**：[arXiv cs.LG (ML)](https://arxiv.org/abs/2607.22724)

**详情页**：https://ai.daily.yangsir.net/daily/20260729-T0-06

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*