---
id: 20260718-T0-08
title: "内省微调IFT：训练小型LLM检测并报告自身内部激活扰动"
title_en: "Introspection Fine-Tuning Trains Small LLMs to Report Internal Activation Changes"
url: https://ai.daily.yangsir.net/daily/20260718-T0-08
issue_date: 2026-07-18
publish_date: 2026-07-17T04:00:00.000Z
category: research
source_name: "arXiv cs.CL (NLP)"
source_url: https://arxiv.org/abs/2607.14111
---

# 内省微调IFT：训练小型LLM检测并报告自身内部激活扰动

研究团队提出内省微调（IFT）方法，训练小型语言模型检测并报告其自身内部激活的扰动。该方法通过注入概念向量到模型的残差流中，测试模型能否感知到这些变化。IFT让小模型具备了自我监控能力，为LLM的可解释性和安全性提供了新的审查维度，使模型能主动报告潜在的错误或操纵。

## English Version

**Introspection Fine-Tuning Trains Small LLMs to Report Internal Activation Changes**

Researchers introduce Introspection Fine-Tuning (IFT), training small language models to detect and report perturbations in their own internal activations. By injecting concept vectors into the model's residual stream, the method tests whether the model can perceive these changes. IFT gives small models self-monitoring capabilities, offering a new dimension for LLM interpretability and safety by enabling active error detection.

---

**来源**：[arXiv cs.CL (NLP)](https://arxiv.org/abs/2607.14111)

**详情页**：https://ai.daily.yangsir.net/daily/20260718-T0-08

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*