---
id: 20260805-T0-09
title: "强化学习新问题：奖励模型重塑支持分布，导致后续目标难以学习"
title_en: "On-Policy RL's Hidden Flaw: Verifier Rewards Can Reshape Support, Making Later Goals Unlearnable"
url: https://ai.daily.yangsir.net/daily/20260805-T0-09
issue_date: 2026-08-05
publish_date: 2026-08-04T04:00:00.000Z
category: research
source_name: "arXiv cs.LG (ML)"
source_url: https://arxiv.org/abs/2608.00220
---

# 强化学习新问题：奖励模型重塑支持分布，导致后续目标难以学习

一篇arXiv论文指出，带可验证奖励的在线强化学习（RLVR）在优化当前目标的同时，可能会使后续目标的成功行为变得过于稀疏而难以被采样和强化。研究者将这种现象称为“验证器诱导的支持重塑”。这意味着训练AI时，为短期目标优化可能会损害其学习长期复杂任务的能力，对RL训练策略的设计有重要启示。

## English Version

**On-Policy RL's Hidden Flaw: Verifier Rewards Can Reshape Support, Making Later Goals Unlearnable**

A new arXiv paper reveals that on-policy reinforcement learning with verifiable rewards (RLVR) can, while improving the current objective, make successful behaviors for subsequent objectives too rare to sample and reinforce. The researchers call this 'verifier-induced support reshaping'. This implies that optimizing for short-term goals during AI training could harm its ability to learn long-term, complex tasks, offering crucial insights for RL training strategy design.

---

**来源**：[arXiv cs.LG (ML)](https://arxiv.org/abs/2608.00220)

**详情页**：https://ai.daily.yangsir.net/daily/20260805-T0-09

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*