---
id: 20260913-T0-02
title: "Real-SWE：用私有企业代码库测AI编程能力，85分登上Hacker News"
title_en: "Real-SWE Benchmarks AI Models on Private Enterprise Codebases"
url: https://ai.daily.yangsir.net/daily/20260913-T0-02
issue_date: 2026-09-13
publish_date: 2026-09-12T20:25:48.000Z
category: research
source_name: "HN AI 精选"
source_url: https://withspecific.com/benchmarks/real-swe
---

# Real-SWE：用私有企业代码库测AI编程能力，85分登上Hacker News

Specific公司推出Real-SWE基准，用私有的真实企业代码库来评测AI模型的软件工程能力。和公开数据集不同，这些代码库不对外可见，模型无法通过记忆训练数据来取巧。该基准在Hacker News上获得85分、54条评论，讨论集中在评测方法是否公平、企业代码隐私如何保护，以及现有SWE-bench类基准是否已经饱和。对关注AI编程工具实际落地效果的人来说，这是一个更接近真实工作场景的评测信号。

## English Version

**Real-SWE Benchmarks AI Models on Private Enterprise Codebases**

Specific has launched Real-SWE, a benchmark that evaluates AI models' software engineering capabilities on private, real-world enterprise codebases. Unlike public datasets, these codebases are not visible externally, preventing models from gaming the benchmark through training data memorization. The benchmark scored 85 points and 54 comments on Hacker News, with discussion focused on evaluation fairness, enterprise code privacy, and whether existing SWE-bench-style benchmarks have saturated. It offers a signal closer to real-world work scenarios for those tracking AI coding tools.

---

**来源**：[HN AI 精选](https://withspecific.com/benchmarks/real-swe)

**详情页**：https://ai.daily.yangsir.net/daily/20260913-T0-02

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*