---
id: 20260825-T0-04
title: "当干净数据有害：研究发现非独立分布数据破坏学习器的理论前提"
title_en: "When Clean Data Hurts: New Work on Learning Under Monotone Corruptions"
url: https://ai.daily.yangsir.net/daily/20260825-T0-04
issue_date: 2026-08-25
publish_date: 2026-08-24T04:00:00.000Z
category: research
source_name: "arXiv cs.LG (ML)"
source_url: https://arxiv.org/abs/2608.20480
---

# 当干净数据有害：研究发现非独立分布数据破坏学习器的理论前提

经典PAC学习理论假设数据独立同分布，一项新研究揭示，当训练集中混入来自无关甚至敌对源、但标签正确的样本时，这种“干净”但非独立同分布的数据会破坏最优学习器的性能。论文将问题形式化为单调腐蚀下的学习，并分析其对模型泛化能力的影响。该结果提醒研究者，在数据收集和清洗时需关注样本分布，单纯追求标签正确并不足够。

## English Version

**When Clean Data Hurts: New Work on Learning Under Monotone Corruptions**

This paper investigates a critical flaw in the PAC learning model: the impact of monotone corruptions where correctly labeled samples from an unrelated or adversarial source are added to an i.i.d. training set. It demonstrates that such 'clean' yet non-i.i.d. data harms the performance of optimal learners that rely on the classic assumption. The study formalizes learning under monotone corruptions and analyzes its impact on generalization, warning that simply ensuring correct labels in datasets is insufficient.

---

**来源**：[arXiv cs.LG (ML)](https://arxiv.org/abs/2608.20480)

**详情页**：https://ai.daily.yangsir.net/daily/20260825-T0-04

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*