从零搭建 vLLM:系统工程师的 LLM 推理课

Build a vLLM from Scratch: An LLM Inference Course for Systems Engineers

Tech-Experiment #LLM#推理引擎#Apple Silicon#开源课程#系统工程
🇨🇳 中文

同一个作者写了 mini-lsm(从零搭存储引擎),现在出了 tiny-llm:从零搭 LLM 推理引擎,4 周课程,目标是让系统工程师真正懂 vLLM 在做什么。

作者是 Alex Chi Z(skyzh),Databricks 工程师,CMU 数据库组出身。tiny-llm 是 mini-lsm 思路在 LLM 推理方向的延伸:不调高层 API,手写每一层。

Apache-2.0,4.8K stars。GitHub: https://github.com/skyzh/tiny-llm


四周课表

Week 1 — Transformer 核心组件

从最基础的部件开始:注意力机制、RoPE 位置编码、Grouped-Query Attention(GQA)、RMSNorm、MLP,然后接上模型加载、解码、采样。

这一周的目标是:不借助 HuggingFace Transformers,把一个可以跑的 Qwen3-4B 推理路径搭起来。

Week 2 — 优化内核

在基础路径跑通之后,开始做性能相关的部分:

  • KV Cache:把历史 token 的 Key/Value 缓存起来,不重复计算
  • 4-bit 量化矩阵向量积:权重量化到 4-bit,矩阵向量乘法用量化内核跑
  • SIMD 矩阵预填充:prefill 阶段用向量指令加速
  • 融合算子:RMSNorm、RoPE、SwiGLU 合并成单个内核,减少内存 round-trip
  • 分块预填充注意力(Tiled Prefill Attention)

Week 3 — 服务基础设施(tiny vLLM)

这是课程的核心:把单次推理变成可以同时处理多个请求的服务:

  • 连续批处理(Continuous Batching):不等批次填满就开始处理,来一个发一个
  • 分块预填充(Chunked Prefill):长 prompt 拆成块,和 decode 步骤交错执行
  • Paged KV Cache:KV Cache 不再预分配固定内存,而是按需分页,解决碎片问题
  • Direct Paged Attention + Paged FlashAttention:在 Paged KV Cache 上直接做注意力计算

可选扩展:MoE(混合专家)、投机解码。

Week 4 — 编码 Agent

课程最后一周从推理系统转到 Agent 系统:实现一个完整的编码 Agent 循环,包括:工作区检查、受批准的文件编辑、经验证的命令执行、效果回执(effect receipt)、断点续跑(checkpoint/resume)、上下文压缩、方向修正、结果评估、分叉与选择(fork-and-select)、有界工具结果。

这部分在 LLM 推理课程里比较少见——大多数推理课在服务层就结束了,tiny-llm 把 Agent 执行层也包了进来。


定位:面向系统工程师,不是 ML 研究者

这门课的目标受众不是想调参的 ML 工程师,而是关注内存带宽、内核占用率、KV Cache 增长、请求调度的系统工程师。

作者的背景也是数据库/存储系统(CMU-DB、TiKV、RisingLight),课程设计逻辑更接近”构建系统”而不是”训练模型”。


平台限制

只支持 Apple Silicon Mac——这不是技术限制不足,而是刻意选择。

Apple Silicon 的统一内存架构(CPU 和 GPU 共享同一块内存)在教学上有优势:不需要处理 CPU↔GPU 数据搬运,可以专注于推理逻辑本身。MLX 作为实现 oracle 和性能基线,learners 通过和 MLX 输出对比来验证自己的实现。

模型固定为 Qwen3-4B,选择理由:GQA、QK normalization、BF16 激活、4-bit 权重——刚好能覆盖课程涉及的所有技术点。


使用方式

# 要求 Python 3.10–3.12,Apple Silicon Mac
pdm install -v
pdm run check-installation
pdm run test-refsol -- -- -k week_1   # 运行 Week 1 参考实现测试

主要依赖:mlx ≥0.32、mlx-lm ≥0.31.3、torch ≥2.6、torchtune、torchao、numpy ≥2.2、nanobind。

仓库结构:

  • tiny_llm/:学习者的实现空间(写你的代码)
  • tiny_llm_ref/:参考答案(测试和 benchmark 用)

和 vLLM 本身的关系

tiny-llm 不是 vLLM 的学习路径,而是vLLM 的教学重现。它把 vLLM 中最关键的几个设计(Paged KV Cache、Continuous Batching、Chunked Prefill)从生产代码里抽出来,在一个可以一步步构建的教学框架里重新实现,帮助工程师真正理解”这些优化在做什么”——而不只是会调 API。


一句话说清楚

tiny-llm 是面向系统工程师的 LLM 推理实战课:4 周手写 Transformer 组件、KV Cache 内核、连续批处理服务层、编码 Agent 循环。只跑 Apple Silicon,用 Qwen3-4B,Apache-2.0。


Apache-2.0。skyzh (Alex Chi Z),Databricks/CMU-DB,4.8K stars,2025-04-19 创建。开源仅供学习参考。


🇬🇧 English

Build a vLLM from Scratch: An LLM Inference Course for Systems Engineers

The same author who wrote mini-lsm (build a storage engine from scratch) is back with tiny-llm: build an LLM inference engine from scratch. Four weeks, Apple Silicon only.

Author: Alex Chi Z (skyzh), engineer at Databricks, formerly at CMU Database Group. tiny-llm applies the mini-lsm philosophy to LLM inference — no high-level API shortcuts, write every layer yourself.

Apache-2.0, 4.8K stars. GitHub: https://github.com/skyzh/tiny-llm


4-Week Curriculum

Week 1 — Core transformer components

Attention, RoPE positional encoding, Grouped-Query Attention (GQA), RMSNorm, MLP, model loading, decoding, sampling. Goal: run Qwen3-4B inference without calling HuggingFace Transformers.

Week 2 — Optimization kernels

  • KV cache: cache past Key/Value tensors, avoid re-computation
  • 4-bit quantized matrix-vector products
  • SIMD matrix prefill
  • Fused RMSNorm/RoPE/SwiGLU kernels (reducing memory round-trips)
  • Tiled prefill attention

Week 3 — Serving infrastructure (the tiny vLLM)

  • Continuous batching: process requests as they arrive instead of waiting for a full batch
  • Chunked prefill: split long prompts into chunks, interleaved with decode steps
  • Paged KV cache: on-demand paging instead of fixed pre-allocation — eliminates fragmentation
  • Direct paged attention + paged FlashAttention

Optional extensions: MoE, speculative decoding.

Week 4 — Coding agent

Full coding agent loop: workspace inspection, approved edits, validated commands, effect receipts, checkpoint/resume, context compaction, steering, outcome evaluation, fork-and-select, bounded tool-result evidence.


Why Systems Engineers, Not ML Researchers

The course is explicitly framed around memory bandwidth, kernel occupancy, KV cache growth, and request scheduling — the same concerns you’d have building a database or storage system. Alex Chi Z’s background (CMU-DB, TiKV, RisingLight) shows in the curriculum design.


Apple Silicon Only

Not a limitation — a deliberate choice. Unified memory means no CPU↔GPU data movement to handle, letting learners focus on inference logic. MLX serves as both the correctness oracle and the performance baseline.

Model: Qwen3-4B — chosen for GQA, QK normalization, BF16 activations, and 4-bit weights that cover every technique in the course.


Setup

# Python 3.10–3.12, Apple Silicon Mac required
pdm install -v
pdm run check-installation
pdm run test-refsol -- -- -k week_1

Key deps: mlx ≥0.32, mlx-lm ≥0.31.3, torch ≥2.6, torchtune, torchao, nanobind.

Repo structure:

  • tiny_llm/: your implementation (exercises go here)
  • tiny_llm_ref/: reference solution (used by tests)

TL;DR

tiny-llm is a hands-on LLM inference course for systems engineers: 4 weeks building transformer components, KV cache kernels, a continuous-batching serving layer, and a coding agent loop — from scratch, Apple Silicon only, Qwen3-4B, Apache-2.0.


Apache-2.0. skyzh (Alex Chi Z), Databricks/CMU-DB, 4.8K stars, created 2025-04-19. For reference only.

💬 评论与讨论

使用 GitHub 账号登录后发表评论

关于本站 · 免责声明

🍄 Mushroom Research Blog 是非营利、免费公开的个人科技观察博客与公众号 XStack18,不接受商业合作、不代表任何企业或机构立场,也不谋求商业利益。我们以个人视角客观中立地记录和分析 AI、Web3 等领域的最新模型发布与技术动态——不止转述新闻标题或二手信息,而是给出有独立思考的深入分析,希望帮更多人获得有价值的一手科技认知。

⚠️ 文中介绍的开源代码与模型,仅供学习交流与技术借鉴。它们大多仍处于早期阶段,有待进一步研究和验证,请勿直接用于工作或生产环境;如需采用,请先自行充分测试,并核实其许可证与安全性。
Open-source code and models featured here are shared for learning and reference only. Most are early-stage and still need further study and verification — please don't use them directly in your work or in production. Test them thoroughly and check their licenses and security first.

  1. 本站文章均为作者基于公开信息的个人研究与观点整理,不代表文中提及的任何公司、产品、模型的官方立场,未与其构成商业关联或合作关系。
  2. 科技行业信息更新极快,我们尽力保证内容准确、及时,但不对完整性、实时性做绝对保证,具体请以相关企业/项目官方公告为准。
  3. 文中引用的第三方商标、产品名称、图片、数据等版权归原权利人所有,我们会尽量注明来源;如你认为存在版权疑问或侵权,请通过下方邮箱联系我们,收到通知后会尽快核实处理(更正、加注来源或删除)。
  4. 文章内容仅为技术科普与个人观点,不构成投资、法律或其他专业建议,据此进行任何决策的后果需自行判断和承担。

📮 侵权 / 勘误 / 合作咨询:[email protected]