---
id: 20260912-T0-08
title: "Qiushi Engine用1000万词训练BabyLM，实现数据高效语言建模新突破"
title_en: "Qiushi Engine Trains BabyLM on 10M Words for Data-Efficient Language Modeling"
url: https://ai.daily.yangsir.net/daily/20260912-T0-08
issue_date: 2026-09-12
publish_date: 2026-09-11T04:00:00.000Z
category: research
source_name: "arXiv cs.CL (NLP)"
source_url: https://arxiv.org/abs/2609.10702
---

# Qiushi Engine用1000万词训练BabyLM，实现数据高效语言建模新突破

Qiushi Engine在BabyLM 2026 Strict-Small赛道开展了一项长期、端到端的自主研究计划，在仅1000万词的语料规模内训练语言模型。研究聚焦于让模型从有限文本中学会利用上下文、泛化到新输入并保留有用能力。该工作展示了从前沿探索到原则指导的模型改进路径，为低资源场景下的语言模型训练提供了可复现的实践参考。对研究者而言，这意味着在极小数据预算下也能系统性地推进模型能力边界。

## English Version

**Qiushi Engine Trains BabyLM on 10M Words for Data-Efficient Language Modeling**

Qiushi Engine conducted a long-horizon, end-to-end autonomous research program on BabyLM 2026 Strict-Small, training language models within a 10-million-word corpus. The work focuses on enabling models to leverage context, generalize to new inputs, and retain useful capabilities from limited text. It maps a path from frontier exploration to principle-guided model improvement, offering a reproducible reference for low-resource language modeling. For researchers, this shows that systematic progress on model capabilities is achievable even under extremely small data budgets.

---

**来源**：[arXiv cs.CL (NLP)](https://arxiv.org/abs/2609.10702)

**详情页**：https://ai.daily.yangsir.net/daily/20260912-T0-08

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*