---
id: 20260826-T0-06
title: "LLM排行榜不中立：选项顺序和提示措辞就能左右排名"
title_en: "No Neutral Harness: LLM Leaderboards Driven by Config-Fragile Items"
url: https://ai.daily.yangsir.net/daily/20260826-T0-06
issue_date: 2026-08-26
publish_date: 2026-08-25T04:00:00.000Z
category: research
source_name: "arXiv cs.AI"
source_url: https://arxiv.org/abs/2608.21382
---

# LLM排行榜不中立：选项顺序和提示措辞就能左右排名

一项研究发现，LLM排行榜的排名高度依赖测试配置，包括选项顺序、提示措辞、答案提取方式等。同一模型在不同配置下得分波动明显，部分题目甚至会对配置变化“脆弱”到改变排名。作者呼吁社区重视榜单的“harness”效应，并建议在评测时公开完整配置以增强可复现性和公平性。

## English Version

**No Neutral Harness: LLM Leaderboards Driven by Config-Fragile Items**

A new paper shows that LLM leaderboards are significantly influenced by harness configuration—option ordering, prompt wording, and answer extraction methods. Model rankings can shift substantially under different configs, with certain items extremely sensitive to setup changes. The authors urge the community to standardize and transparently document harness settings to improve the reproducibility and fairness of LLM evaluations.

---

**来源**：[arXiv cs.AI](https://arxiv.org/abs/2608.21382)

**详情页**：https://ai.daily.yangsir.net/daily/20260826-T0-06

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*