---
id: 20260915-T0-06
title: "GAUGE：LLM评委在用户模拟评估中何时不可信"
title_en: "GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Agent Evaluation"
url: https://ai.daily.yangsir.net/daily/20260915-T0-06
issue_date: 2026-09-15
publish_date: 2026-09-14T04:00:00.000Z
category: research
source_name: "arXiv cs.CL (NLP)"
source_url: https://arxiv.org/abs/2609.12191
---

# GAUGE：LLM评委在用户模拟评估中何时不可信

arXiv 新论文 GAUGE 针对任务型 LLM 智能体的低成本离线评估流程提出质疑。常见做法是：用 persona 驱动的 LLM 用户模拟器与候选智能体对话，再由 LLM 评委给对话记录打分，分数高者胜出。论文指出这套流程存在系统性偏差，会选出实际表现并非最优的智能体。研究给出了判断 LLM 评委何时不可信的方法，为团队在选择评估方案时提供参考。对依赖离线评测做模型选型的开发者来说，这意味着现有的自动评估结果需要重新审视。

## English Version

**GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Agent Evaluation**

A new arXiv paper, GAUGE, challenges the low-cost offline evaluation pipeline widely used for task-oriented LLM agents. The common setup has persona-driven LLM user-simulators converse with each candidate agent, an LLM-as-a-judge scores the transcripts, and the highest scorer wins. The paper identifies systematic biases in this pipeline that can promote agents that are not actually the best performers. It offers a method for determining when LLM judges should not be trusted, giving teams a way to sanity-check their evaluation choices. For developers relying on offline evals for model selection, existing automated ranking results may need re-examination.

---

**来源**：[arXiv cs.CL (NLP)](https://arxiv.org/abs/2609.12191)

**详情页**：https://ai.daily.yangsir.net/daily/20260915-T0-06

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*