---
id: 20260709-T0-14
title: "OpenAI分析发现编程基准SWE-Bench Pro存在可靠性问题"
title_en: "OpenAI Analysis Flags Reliability Issues in Coding Benchmark SWE-Bench Pro"
url: https://ai.daily.yangsir.net/daily/20260709-T0-14
issue_date: 2026-07-09
publish_date: 2026-07-08T13:00:00.000Z
category: research
source_name: "OpenAI News"
source_url: https://openai.com/index/separating-signal-from-noise-coding-evaluations
---

# OpenAI分析发现编程基准SWE-Bench Pro存在可靠性问题

OpenAI 发布分析报告，指出流行的 AI 编码能力基准测试 SWE-Bench Pro 存在信号噪声问题。报告显示，该基准可能无法准确反映模型在真实编程任务中的表现，部分评估指标存在偏差和可重复性差的问题。这一发现引发对当前 AI 编程评估方法可靠性的讨论。

## English Version

**OpenAI Analysis Flags Reliability Issues in Coding Benchmark SWE-Bench Pro**

OpenAI published an analysis revealing issues with SWE-Bench Pro, a popular coding benchmark for AI models. The report indicates reliability and accuracy concerns, suggesting the benchmark may not accurately reflect real-world coding performance and has reproducibility issues. This raises questions about the validity of current AI coding evaluation methods.

---

**来源**：[OpenAI News](https://openai.com/index/separating-signal-from-noise-coding-evaluations)

**详情页**：https://ai.daily.yangsir.net/daily/20260709-T0-14

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*