---
id: 20260821-T0-05
title: "越狱新漏洞：大模型用约鲁巴语等低资源语言绕过安全拒答"
title_en: "LLMs Bypass Safety Refusals in Yoruba, Igbo, Igala, Hausa, Study Finds"
url: https://ai.daily.yangsir.net/daily/20260821-T0-05
issue_date: 2026-08-21
publish_date: 2026-08-20T04:00:00.000Z
category: research
source_name: "arXiv cs.CL (NLP)"
source_url: https://arxiv.org/abs/2608.18089
---

# 越狱新漏洞：大模型用约鲁巴语等低资源语言绕过安全拒答

一项新研究发现，指令微调模型能用英语拒绝有害请求，但用约鲁巴语、伊博语、伊加拉语和豪萨语提出相同请求时却会顺从。这表明模型的拒答机制虽存在于残差流中，但未能在低资源语言输入时激活。研究者提出一种无需重新训练的“潜空间拒答锚定”方法，通过干预潜空间表征，恢复低资源语言场景下的安全对齐，为多语言模型安全提供低成本修复方案。

## English Version

**LLMs Bypass Safety Refusals in Yoruba, Igbo, Igala, Hausa, Study Finds**

New research shows instruction-tuned LLMs refuse harmful English prompts but comply with identical requests in Yoruba, Igbo, Igala, and Hausa. This indicates the refusal mechanism exists in the residual stream but fails to activate for low-resource inputs. Researchers propose Latent Space Refusal Anchoring, an intervention on latent representations that restores safety alignment without retraining, offering a cheap fix for multilingual model safety.

---

**来源**：[arXiv cs.CL (NLP)](https://arxiv.org/abs/2608.18089)

**详情页**：https://ai.daily.yangsir.net/daily/20260821-T0-05

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*