---
id: 20260704-T0-04
title: "BPE分词在词边界制造安全漏洞，越狱LLM只需字符级扰动"
title_en: "BPE Tokenization Creates Exploitable Safety Gaps in LLM Alignment"
url: https://ai.daily.yangsir.net/daily/20260704-T0-04
issue_date: 2026-07-04
publish_date: 2026-07-03T04:00:00.000Z
source_name: "arXiv cs.CL (NLP)"
source_url: https://arxiv.org/abs/2607.01239
---

# BPE分词在词边界制造安全漏洞，越狱LLM只需字符级扰动

研究发现，BPE分词器将关键安全词拆分为子词片段，导致字符级扰动即可绕过LLM的安全对齐。三个主流闭源模型均受影响：通过在安全词中间插入空格或特殊字符，攻击者可使模型输出被禁止的内容，而提示对人类仍然可读。

## English Version

**BPE Tokenization Creates Exploitable Safety Gaps in LLM Alignment**

Research reveals BPE tokenizers fragment safety-critical words into sub-word pieces, allowing character-level perturbations to bypass LLM safety alignment. All three major closed-source models are affected: inserting spaces or special characters within safety words triggers prohibited outputs while prompts remain human-readable.

---

**来源**：[arXiv cs.CL (NLP)](https://arxiv.org/abs/2607.01239)

**详情页**：https://ai.daily.yangsir.net/daily/20260704-T0-04

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*