---
id: 20260624-T0-08
title: "VeriBound：利用形式化验证提升过程奖励模型泛化性"
title_en: "VeriBound: Improving PRM Generalization with Formal Verification Tools"
url: https://ai.daily.yangsir.net/daily/20260624-T0-08
issue_date: 2026-06-24
publish_date: 2026-06-23T04:00:00.000Z
category: research
source_name: "arXiv cs.CL (NLP)"
source_url: https://arxiv.org/abs/2606.20740
---

# VeriBound：利用形式化验证提升过程奖励模型泛化性

针对过程奖励模型（PRM）训练数据获取难、噪声大的问题，新研究提出了 VeriBound 方法。该方案利用形式化验证工具合成高质量训练数据，并基于 PAC-Bayesian 理论给出了泛化误差界。实验证明，该方法能有效提升 PRM 在复杂推理任务中的准确性和稳定性。

## English Version

**VeriBound: Improving PRM Generalization with Formal Verification Tools**

To address the bottleneck of noisy and costly training data for Process Reward Models (PRMs), researchers propose VeriBound. This method uses formal verification tools to synthesize high-quality data and provides generalization bounds based on PAC-Bayesian theory, significantly improving PRM accuracy in complex reasoning tasks.

---

**来源**：[arXiv cs.CL (NLP)](https://arxiv.org/abs/2606.20740)

**详情页**：https://ai.daily.yangsir.net/daily/20260624-T0-08

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*