---
id: 20260709-T0-10
title: "离线强化学习学会控制LLM Agent的执行框架"
title_en: "Learning to Control LLM Agent Harnesses with Offline Reinforcement Learning"
url: https://ai.daily.yangsir.net/daily/20260709-T0-10
issue_date: 2026-07-09
publish_date: 2026-07-08T04:00:00.000Z
category: research
source_name: "arXiv cs.LG (ML)"
source_url: https://arxiv.org/abs/2607.05458
---

# 离线强化学习学会控制LLM Agent的执行框架

该研究提出将 LLM Agent 的执行框架（harness）视为可学习的控制层，而非固定基础设施。传统方法通过修改提示词或模型来改进 Agent，而新方法使用离线强化学习优化 Agent 如何调用工具、管理状态和执行步骤。实验表明，学习控制执行流程能比仅优化 LLM 本身带来更显著的 Agent 性能提升。

## English Version

**Learning to Control LLM Agent Harnesses with Offline Reinforcement Learning**

This research proposes treating the LLM agent execution harness as a learnable control layer, not fixed infrastructure. Instead of just tweaking prompts or models, it uses offline reinforcement learning to optimize how the agent invokes tools, manages state, and executes steps. Results show that learning to control the harness yields greater agent performance gains than optimizing the LLM alone.

---

**来源**：[arXiv cs.LG (ML)](https://arxiv.org/abs/2607.05458)

**详情页**：https://ai.daily.yangsir.net/daily/20260709-T0-10

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*