---
id: 20260814-T0-10
title: "PAIR方法优化RLVR：自适应分配预算，降低推理成本"
title_en: "PAIR: Adaptive Rollout Allocation Cuts RLVR Compute Costs"
url: https://ai.daily.yangsir.net/daily/20260814-T0-10
issue_date: 2026-08-14
publish_date: 2026-08-13T04:00:00.000Z
category: research
source_name: "arXiv cs.LG (ML)"
source_url: https://arxiv.org/abs/2608.11368
---

# PAIR方法优化RLVR：自适应分配预算，降低推理成本

arXiv新论文提出PAIR（Pairwise-Aware Inclusion Reweighting）方法，用于强化学习可验证奖励（RLVR）场景下的自适应rollout分配。RLVR大部分算力用于生成长推理轨迹，现有分配器按点级指标分配预算，效率有限。PAIR通过成对感知的包含重加权，优化提示、轨迹和token的预算分配，在保持性能的同时显著降低训练成本，推动RLVR在复杂推理任务中的落地。

## English Version

**PAIR: Adaptive Rollout Allocation Cuts RLVR Compute Costs**

An arXiv paper introduces PAIR (Pairwise-Aware Inclusion Reweighting), an adaptive rollout allocation method for reinforcement learning with verifiable rewards (RLVR). Since RLVR spends most compute on long reasoning trajectories, PAIR optimizes budget allocation across prompts, rollouts, and tokens via pairwise-aware reweighting, slashing training costs while sustaining performance, advancing RLVR adoption in complex reasoning tasks.

---

**来源**：[arXiv cs.LG (ML)](https://arxiv.org/abs/2608.11368)

**详情页**：https://ai.daily.yangsir.net/daily/20260814-T0-10

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*