---
id: 20260811-T0-11
title: "平均奖励MDP稳健学习：样本复杂度达极小极大最优"
title_en: "Average-Reward MDPs: Minimax-Optimal Robust Learning via Plug-in Reductions"
url: https://ai.daily.yangsir.net/daily/20260811-T0-11
issue_date: 2026-08-11
publish_date: 2026-08-10T04:00:00.000Z
source_name: "arXiv cs.LG (ML)"
source_url: https://arxiv.org/abs/2608.06545
---

# 平均奖励MDP稳健学习：样本复杂度达极小极大最优

一篇新论文研究了稳健平均奖励马尔可夫决策过程（MDP）的样本复杂度，提出了一种基于插件归约（plug-in reduction）的方法，在模型不确定下实现极小极大最优的学习效率。该研究为分布鲁棒MDP提供了理论保障，明确了在ε误差下需要多少样本才能学到最优稳健策略。这项成果为在医疗、金融等高风险动态决策场景中，如何更少样本地应对环境不确定性提供了理论指引。

## English Version

**Average-Reward MDPs: Minimax-Optimal Robust Learning via Plug-in Reductions**

This paper investigates sample complexity in robust average-reward Markov decision processes (MDPs). It introduces a plug-in reduction method achieving minimax-optimal learning efficiency under model uncertainty. Theoretical bounds clarify the necessary and sufficient samples for an ε-optimal robust policy. This provides practical guidance for high-stakes dynamic decisions, such as healthcare or finance, where data is limited and model misspecification is common.

---

**来源**：[arXiv cs.LG (ML)](https://arxiv.org/abs/2608.06545)

**详情页**：https://ai.daily.yangsir.net/daily/20260811-T0-11

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*