---
id: 20260906-T0-07
title: "新基准DuplexSpeechBench-IFEval：评估全双工语音代理的隐式指令遵循能力"
title_en: "New Benchmark DuplexSpeechBench-IFEval Tests Implicit Instruction Following"
url: https://ai.daily.yangsir.net/daily/20260906-T0-07
issue_date: 2026-09-06
publish_date: 2026-09-05T04:00:00.000Z
category: research
source_name: "arXiv cs.AI"
source_url: https://arxiv.org/abs/2609.03423
---

# 新基准DuplexSpeechBench-IFEval：评估全双工语音代理的隐式指令遵循能力

全双工语音代理需持续决策何时聆听、插话或处理语音重叠。现有基准主要测试显式指令遵循，但实际部署中更多依赖隐式理解。新提出的DuplexSpeechBench-IFEval基准填补了这一空白，专门评估语音代理在复杂对话场景下理解并执行未明确言明的指令的能力。这为开发者提供了更贴近真实应用的测试工具，有助于改进语音助手的交互自然度和鲁棒性。

## English Version

**New Benchmark DuplexSpeechBench-IFEval Tests Implicit Instruction Following**

Full-duplex voice agents must continuously decide when to listen, interrupt, or handle overlaps. Existing benchmarks focus on explicit instructions, but real deployments rely on implicit understanding. The new DuplexSpeechBench-IFEval benchmark fills this gap by evaluating agents' ability to follow unspoken instructions in complex scenarios. It offers developers a more realistic testing tool to improve the naturalness and robustness of voice assistants.

---

**来源**：[arXiv cs.AI](https://arxiv.org/abs/2609.03423)

**详情页**：https://ai.daily.yangsir.net/daily/20260906-T0-07

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*