---
id: 20260916-T0-02
title: "同一患者、不同医嘱：临床LLM智能体重复运行可靠性存疑"
title_en: "Same Patient, Different Orders: Clinical LLM Agents Show Low Action-Level Reliability Across Repeated Runs"
url: https://ai.daily.yangsir.net/daily/20260916-T0-02
issue_date: 2026-09-16
publish_date: 2026-09-15T04:00:00.000Z
category: research
source_name: "arXiv cs.CL (NLP)"
source_url: https://arxiv.org/abs/2609.13582
---

# 同一患者、不同医嘱：临床LLM智能体重复运行可靠性存疑

一项新研究揭示了临床LLM智能体的可靠性问题：在相同输入下，智能体可能给出相同诊断结论，但每次运行实际开出的检查、处方和转诊医嘱却存在实质性差异。目前临床智能体基准测试通常每个任务只运行一次，无法捕捉这种行动层面的不一致性。研究呼吁建立多轮重复运行的动作级评估标准，否则智能体在真实医疗场景中的安全性难以保证。

## English Version

**Same Patient, Different Orders: Clinical LLM Agents Show Low Action-Level Reliability Across Repeated Runs**

A new study reveals reliability issues in clinical LLM agents: given identical inputs, an agent may produce the same diagnostic verdict while filing materially different test orders, prescriptions, and referrals on each run. Current clinical agent benchmarks typically score one run per task, failing to capture this action-level inconsistency. The research calls for multi-run, action-level evaluation standards, warning that agent safety in real clinical settings cannot be guaranteed otherwise.

---

**来源**：[arXiv cs.CL (NLP)](https://arxiv.org/abs/2609.13582)

**详情页**：https://ai.daily.yangsir.net/daily/20260916-T0-02

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*