---
id: 20260710-T0-09
title: "AgentLens：用生产环境数据评估编程智能体的新基准"
title_en: "AgentLens: New Benchmark Evaluates Coding Agents with Production Trajectory Data"
url: https://ai.daily.yangsir.net/daily/20260710-T0-09
issue_date: 2026-07-10
publish_date: 2026-07-09T04:00:00.000Z
category: research
source_name: "arXiv cs.AI"
source_url: https://arxiv.org/abs/2607.06624
---

# AgentLens：用生产环境数据评估编程智能体的新基准

研究人员提出AgentLens，一个基于生产环境评估的交互式编程智能体基准。传统评测只记录任务是否通过，但AgentLens关注整个执行轨迹（包括中间步骤和错误），更贴近真实用户对智能体的使用体验。

## English Version

**AgentLens: New Benchmark Evaluates Coding Agents with Production Trajectory Data**

Researchers introduced AgentLens, a production-assessed benchmark for interactive code agents. Unlike traditional single-pass/fail metrics, it evaluates the entire execution trajectory, capturing intermediate steps and errors for a more realistic user experience assessment.

---

**来源**：[arXiv cs.AI](https://arxiv.org/abs/2607.06624)

**详情页**：https://ai.daily.yangsir.net/daily/20260710-T0-09

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*