---
id: 20260718-T0-04
title: "BPO：一种沙盒原生语言智能体强化学习新方法"
title_en: "BPO: Sandbox-native reinforcement learning for LLM agents outperforms PPO and RLOO"
url: https://ai.daily.yangsir.net/daily/20260718-T0-04
issue_date: 2026-07-18
publish_date: 2026-07-17T04:00:00.000Z
category: research
source_name: "arXiv cs.LG (ML)"
source_url: https://arxiv.org/abs/2607.14171
---

# BPO：一种沙盒原生语言智能体强化学习新方法

新论文提出分支策略优化（Branching Policy Optimization, BPO），一种专为沙盒环境设计的语言智能体强化学习方法。与PPO、RLOO、GRPO等依赖RLHF的rollout拓扑不同，BPO利用沙盒原生的分支特性，允许智能体在环境中并行探索多条路径，从而提升训练效率和最终性能。

## English Version

**BPO: Sandbox-native reinforcement learning for LLM agents outperforms PPO and RLOO**

A new paper introduces Branching Policy Optimization (BPO), a reinforcement learning method designed for sandbox environments. Unlike PPO, RLOO, and GRPO which inherit RLHF rollout topology, BPO leverages sandbox-native branching to let agents explore multiple paths in parallel, improving training efficiency and final performance.

---

**来源**：[arXiv cs.LG (ML)](https://arxiv.org/abs/2607.14171)

**详情页**：https://ai.daily.yangsir.net/daily/20260718-T0-04

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*