---
id: 20260715-T0-06
title: "MawForge：让MoE大模型在本地设备上高效运行的推理框架"
title_en: "MawForge: Memory-Bounded Expert Materialization for Local MoE Inference"
url: https://ai.daily.yangsir.net/daily/20260715-T0-06
issue_date: 2026-07-15
publish_date: 2026-07-14T04:00:00.000Z
source_name: "arXiv cs.LG (ML)"
source_url: https://arxiv.org/abs/2607.09686
---

# MawForge：让MoE大模型在本地设备上高效运行的推理框架

arXiv新论文提出MawForge框架，专为在资源受限的本地设备上运行稀疏混合专家（MoE）大模型而设计。该方法通过内存受限的专家实例化，将全模型加载需求降至最低，同时保留推理性能。这使得在个人电脑上运行大参数MoE模型成为可能。

## English Version

**MawForge: Memory-Bounded Expert Materialization for Local MoE Inference**

A new arXiv paper introduces MawForge, a framework for running sparse Mixture-of-Experts (MoE) LLMs on resource-constrained local devices. It minimizes full model loading via memory-bounded expert materialization while preserving inference performance, enabling large MoE models on personal computers.

---

**来源**：[arXiv cs.LG (ML)](https://arxiv.org/abs/2607.09686)

**详情页**：https://ai.daily.yangsir.net/daily/20260715-T0-06

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*