---
id: 20260626-T0-15
title: "利用中间层熵值动态检测大模型越狱攻击"
title_en: "Detecting Jailbreaks from Entropy Dynamics in Intermediate Layers"
url: https://ai.daily.yangsir.net/daily/20260626-T0-15
issue_date: 2026-06-26
publish_date: 2026-06-25T04:00:00.000Z
category: research
source_name: "arXiv cs.CL (NLP)"
source_url: https://arxiv.org/abs/2606.25182
---

# 利用中间层熵值动态检测大模型越狱攻击

新研究提出了一种通过监测大语言模型中间层熵值动态来检测越狱攻击的方法。大多数防御机制集中在输入或输出层，而该方法揭示了模型内部在处理恶意提示时的熵值变化，为防御安全训练后的模型漏洞提供了新思路。

## English Version

**Detecting Jailbreaks from Entropy Dynamics in Intermediate Layers**

A new paper proposes detecting jailbreak attacks by analyzing entropy dynamics in the intermediate layers of LLMs. Unlike typical input/output defenses, this method observes internal entropy shifts to identify policy-violating prompts.

---

**来源**：[arXiv cs.CL (NLP)](https://arxiv.org/abs/2606.25182)

**详情页**：https://ai.daily.yangsir.net/daily/20260626-T0-15

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*