---
id: 20260822-T0-18
title: "多模态LLM基准测试：系统消息下的合规、能力与冲突"
title_en: "New Benchmark: MLLMs Under System Messages"
url: https://ai.daily.yangsir.net/daily/20260822-T0-18
issue_date: 2026-08-22
publish_date: 2026-08-21T04:00:00.000Z
category: research
source_name: "arXiv cs.CL (NLP)"
source_url: https://arxiv.org/abs/2608.19207
---

# 多模态LLM基准测试：系统消息下的合规、能力与冲突

一篇新论文针对多模态大语言模型（MLLMs）在系统消息约束下的表现，提出了新的基准测试。目前的基准要么只测试文本约束，要么把约束嵌入用户指令中，忽略了系统消息的影响。该研究专门评估了在系统消息设定下，多模态模型的行为合规性、执行能力以及二者之间的冲突，揭示了现实部署中的潜在问题。

## English Version

**New Benchmark: MLLMs Under System Messages**

A new paper proposes a benchmark for evaluating Multimodal LLMs under system messages. Existing benchmarks either test text-only constraints or embed them in user turns, ignoring system-message impact. This research specifically assesses compliance, capability, and conflict in MLLMs under system instructions, revealing potential issues in production deployments.

---

**来源**：[arXiv cs.CL (NLP)](https://arxiv.org/abs/2608.19207)

**详情页**：https://ai.daily.yangsir.net/daily/20260822-T0-18

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*