---
id: 20260825-T0-06
title: "RLVR奖励模型存在多语言偏见，数学推理训练对非英语不公"
title_en: "Multilingual Verifier Bias Found in RLVR, Hindering Non-English Math Reasoning"
url: https://ai.daily.yangsir.net/daily/20260825-T0-06
issue_date: 2026-08-25
publish_date: 2026-08-24T04:00:00.000Z
category: research
source_name: "arXiv cs.CL (NLP)"
source_url: https://arxiv.org/abs/2608.20362
---

# RLVR奖励模型存在多语言偏见，数学推理训练对非英语不公

研究揭示了强化学习验证奖励（RLVR）中的一个关键问题：当使用回答验证器作为奖励函数训练大模型数学推理时，该验证器被假设为语言无关，但实际上存在跨语言偏见。论文通过基准测试和回滚诊断，确认模型在不同语言上选择正确答案的能力存在瓶颈，即交叉语言选择瓶颈。这导致非英语环境下的训练效果受限。

## English Version

**Multilingual Verifier Bias Found in RLVR, Hindering Non-English Math Reasoning**

A new study challenges the assumption that answer verifiers in Reinforcement Learning with Verifiable Rewards (RLVR) are language-neutral. It reveals a significant multilingual verifier bias, identifying a 'cross-lingual selection bottleneck'. Through benchmark tests and rollout diagnosis, the authors show that models struggle to select correct answers in different languages, leading to suboptimal training outcomes for non-English mathematical reasoning. This highlights a critical flaw in the standard recipe for training LLMs on multilingual reasoning tasks.

---

**来源**：[arXiv cs.CL (NLP)](https://arxiv.org/abs/2608.20362)

**详情页**：https://ai.daily.yangsir.net/daily/20260825-T0-06

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*