---
id: 20260910-T0-09
title: "新研究质疑：仅靠安全监控器无法有效阻止大模型有害输出"
title_en: "Study: Safety Monitors Alone Don't Prevent Model Compliance"
url: https://ai.daily.yangsir.net/daily/20260910-T0-09
issue_date: 2026-09-10
publish_date: 2026-09-09T04:00:00.000Z
category: research
source_name: "arXiv cs.CL (NLP)"
source_url: https://arxiv.org/abs/2609.05797
---

# 新研究质疑：仅靠安全监控器无法有效阻止大模型有害输出

一项来自 arXiv 的研究（编号2609.05797）指出，当前对大语言模型安全监控器的评估方式存在根本缺陷。现有评估依赖“召回率”（recall）来衡量监控器对有害请求的拦截能力，但忽视了关键因素：模型本身是否会遵从该请求。如果模型根本不会执行有害行为，拦截并无必要。研究者提出，监控器的有效性应当结合模型的遵从性来评估，才能真实反映其对实际危害的预防作用。

## English Version

**Study: Safety Monitors Alone Don't Prevent Model Compliance**

A new arXiv study challenges current evaluation methods for safety monitors in large language models by pointing out a fundamental flaw: they are assessed on recall against harmfulness labels without accounting for whether the model would actually comply with the flagged request. If a model refuses harmful behavior anyway, intercepting it is unnecessary. The researchers argue monitors should be evaluated in conjunction with model compliance to measure real-world harm prevention.

---

**来源**：[arXiv cs.CL (NLP)](https://arxiv.org/abs/2609.05797)

**详情页**：https://ai.daily.yangsir.net/daily/20260910-T0-09

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*