---
id: 20260925-T0-02
title: "Agent重复任务表现不一致：同一任务跑三次，38%到74%的结果对不上"
title_en: "Agents Fail Consistency Test: 38%–74% Disagreement Across Repeated Runs"
url: https://ai.daily.yangsir.net/daily/20260925-T0-02
issue_date: 2026-09-25
publish_date: 2026-09-24T04:00:00.000Z
category: research
source_name: "arXiv cs.AI"
source_url: https://arxiv.org/abs/2609.25299
---

# Agent重复任务表现不一致：同一任务跑三次，38%到74%的结果对不上

一项研究对42个任务各运行三次，发现根据模型不同，38%到74%的回答不一致。研究指出，一致性是买家、审计方和监管机构的基本要求，而当前Agent并不具备。论文提出技能应形成习惯，以提升重复任务中的稳定性。这一问题直接影响Agent在需要可重复、可审计结果的场景中的可用性，比如合规审查和自动化运维。

## English Version

**Agents Fail Consistency Test: 38%–74% Disagreement Across Repeated Runs**

A study ran 42 tasks three times each and found that, depending on the model, 38% to 74% of answers disagreed. The research notes that consistency is a baseline requirement for buyers, auditors, and regulators, yet current agents lack it. The paper argues skills should form habits to improve stability on repeat tasks. This directly limits agent usability in scenarios that demand reproducible, auditable results, such as compliance reviews and automated operations.

---

**来源**：[arXiv cs.AI](https://arxiv.org/abs/2609.25299)

**详情页**：https://ai.daily.yangsir.net/daily/20260925-T0-02

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*