---
id: 20260717-T0-13
title: "GFlowRL：用分布匹配RL提升大模型推理多样性"
title_en: "GFlowRL: Scaling Distribution-Matching RL to Large Language Models"
url: https://ai.daily.yangsir.net/daily/20260717-T0-13
issue_date: 2026-07-17
publish_date: 2026-07-16T04:00:00.000Z
category: research
source_name: "arXiv cs.CL (NLP)"
source_url: https://arxiv.org/abs/2607.13394
---

# GFlowRL：用分布匹配RL提升大模型推理多样性

新论文提出GFlowRL方法，将生成流网络（GFlowNets）应用于大语言模型的强化学习训练。与传统奖励最大化不同，GFlowRL通过匹配奖励分布来鼓励模型生成多样化推理路径，避免坍塌到单一最优解，特别适合大推理模型。

## English Version

**GFlowRL: Scaling Distribution-Matching RL to Large Language Models**

A new paper introduces GFlowRL, applying Generative Flow Networks (GFlowNets) to RL training for LLMs. Instead of maximizing a single reward, GFlowRL matches the reward distribution to encourage diverse reasoning paths, avoiding collapse to a dominant mode, especially useful for large reasoning models.

---

**来源**：[arXiv cs.CL (NLP)](https://arxiv.org/abs/2607.13394)

**详情页**：https://ai.daily.yangsir.net/daily/20260717-T0-13

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*