---
id: 20260925-T0-04
title: "Terminal-Bench任务为何难？研究区分真实难度与伪难度"
title_en: "Why Are Terminal-Bench Tasks Hard? Separating Real from Fake Difficulty"
url: https://ai.daily.yangsir.net/daily/20260925-T0-04
issue_date: 2026-09-25
publish_date: 2026-09-24T04:00:00.000Z
category: research
source_name: "arXiv cs.LG (ML)"
source_url: https://arxiv.org/abs/2609.26826
---

# Terminal-Bench任务为何难？研究区分真实难度与伪难度

前沿基准需要当前模型无法解决的任务，但无人能解的任务不一定真的难。研究指出，同样的零通过率可能来自真实能力缺口，也可能来自缺失上下文或损坏的参考实现。作者构建了一个经过裁决的Agent语料库，用以区分真实难度与伪难度。这一工作有助于基准设计者剔除无效难题，让评测结果更准确反映模型的实际能力。

## English Version

**Why Are Terminal-Bench Tasks Hard? Separating Real from Fake Difficulty**

Frontier benchmarks need tasks current models cannot solve, but a task no model solves is not automatically hard. The research notes the same zero pass rate can come from a real capability gap, or from missing context or a broken reference implementation. The authors built an adjudicated agentic corpus to separate genuine hardness from fake hardness. This helps benchmark designers remove invalid hard tasks so evaluations more accurately reflect real model capability.

---

**来源**：[arXiv cs.LG (ML)](https://arxiv.org/abs/2609.26826)

**详情页**：https://ai.daily.yangsir.net/daily/20260925-T0-04

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*