Strata:12GB 显卡跑 125B 模型,MoE 专家热缓存打破 VRAM 限制

Strata: Run a 125B Model on a 12 GB GPU with MoE Expert Hot Cache

Tech-Experiment #本地推理#MoE#GPU推理#Qwen#开源拆解#推理优化#C++
🇨🇳 中文

开源仅供学习:本文所涉项目均来自公开仓库,分析仅供技术研究。


核心问题

125B 参数的模型,正常来说要跑在多台 A100 或者大内存的服务器 GPU 上。Qwen3.8-Flash-Next 是 125B MoE(Mixture of Experts)架构——参数量是 125B,但每次推理实际只激活很小一部分专家网络。

Strata 针对这个特性专门做了一套推理引擎:把常用的专家缓存在 GPU 上,不常用的放在内存里按需读取。结果是——12GB 显卡 + 64GB 系统内存,能跑这个模型,最快 93 tokens/s。

仓库:github.com/Niko1221/Strata
Stars:5500+ | MIT | C++ | 作者:Niko Veit(Niko1221)


工作原理

MoE 的推理特性

MoE 模型不是每个 token 都走所有参数。Qwen3.8-Flash-Next 的名字里”Flash Next”就暗示了这一点——激活参数量远小于总参数量,每次推理只用到模型权重的一小部分。

Strata 的核心是利用这个特性做异构存储:

  • 热专家留 GPU:被频繁调用的专家(expert layer)缓存在显存里,访问延迟最低
  • 冷专家从 RAM 读:不常用的专家放在系统内存,需要时再加载
  • 自适应驱逐:Strata 用指数衰减算法动态维护缓存,使用越频繁的专家越不容易被驱逐
  • 使用画像持久化:每 10 分钟保存一次当前工作负载的专家使用分布,下次启动直接加载,跳过冷启动阶段

冷专家在 CPU 上用 AVX-512/AVX2 指令集并行计算,和 GPU 上的热专家并行进行,不是纯串行等待。

MTP 投机解码

Strata 在此基础上加了 Multi-Token Prediction(MTP)投机解码。一个小型 draft layer 先猜接下来几个 token,主模型批量验证猜测,正确的直接接受,不对的再走正常生成流程。

效果:短对话场景下 1.6–1.8x 的额外加速,设置为 MTP n-max 4(最多一次猜 4 个 token)。

专用 CUDA 核

Strata 的 CUDA kernel 不是通用的,而是针对 Qwen3.8-Flash-Next 架构定制重写,官方 YouTube 视频标题里称”6x faster than llama.cpp”——这个数字要打折,因为 llama.cpp 本身不是为这类专家 offload 场景优化的,但定制 kernel 确实比通用方案快很多。


性能数据

作者用 RTX 5070(12GB)测的基准数字:

量化方案短对话速度128K 长上下文速度
Q2_0~93 t/s~74 t/s
IQ3_S~53 t/s~46 t/s

用户实测在 RTX 5070 + RTX 3090(24GB)上也有 75-96 t/s 的报告,接近官方数字。

模型文件:ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF,总大小约 75.8GB,全部放在系统内存和磁盘,显存只存当前活跃专家。

需要注意的性能下降:连续对话时,有用户报告(Issue #528)在 RTX 5090 上,持续对话会出现约 4-6x 的速度下降——这是已知问题,v0.1.36 仍在复现调查中。短对话场景不受影响。


安装与使用

一键安装脚本覆盖 Windows 和 Linux,不需要手动配置 CUDA 环境:

# Linux
curl -fsSL https://raw.githubusercontent.com/Niko1221/Strata/main/install.sh | bash

# Windows:下载 install.bat 运行

脚本自动下载模型(75.8GB,首次需要较长时间)、配置推理引擎、启动服务。

启动后提供:

  • 浏览器 UI:本地 Web 界面,直接聊天
  • OpenAI 兼容 API:http://localhost:PORT/v1/chat/completions
  • Anthropic 兼容 API:http://localhost:PORT/v1/messages
  • 可选多模态输入:支持图片(mmproj,8-12GB 卡要注意多模态投影器会多占约 907MB 显存)

硬件门槛实际在哪里

“12GB 显卡”是真的,但完整的要求是:

硬件要求
GPUNVIDIA RTX 30/40/50 系,≥12GB 显存(8GB 有 fork 版本)
系统内存64GB RAM(这才是真正的门槛)
磁盘约 80GB 空闲(模型 + 程序)
OSWindows 或 Linux
后端NVIDIA Only(无 AMD、无 Mac、无 CPU-only)

大多数消费级主机的内存配置是 16-32GB,64GB 需要专门升级。GPU 的 12GB 反而不是最难满足的那个条件。


关键约束

只跑 Qwen3.8-Flash-Next。Strata 不是通用推理引擎,它的 CUDA kernel 和缓存逻辑都是为这一个模型架构定制的。如果想跑其他模型,不适用。

NVIDIA 独占。项目没有 AMD ROCm 或 Metal(Mac)后端。

持续对话有已知性能问题(Issue #528,v0.1.36 仍在调查)。短对话不受影响。

MIT 协议,商用无限制。


关键数字

指标值
Stars5,500+
LicenseMIT
语言C++
运行模型Qwen3.8-Flash-Next(125B MoE)
最低显存12GB NVIDIA
最低系统内存64GB RAM
模型大小~75.8GB
最快速度(Q2_0)~93 t/s(短对话,RTX 5070)
最快速度(IQ3_S)~53 t/s(短对话,RTX 5070)
当前版本v0.1.36
平台Windows + Linux

综合判断

Strata 做的事情在技术上是成立的:MoE 模型天生适合这种异构存储推理方式——每个 token 只激活一小部分专家,所以不需要把全部权重放在 GPU 上。Strata 把这个特性工程化,配合定制 CUDA kernel 和投机解码,把 125B 模型压到了消费级硬件可用的范围。

速度是真实的,96 t/s 对于一个 125B 模型来说确实很快。但”12GB 显卡”这个说法省略了”还需要 64GB 系统内存”这个同样重要的条件——这不是随便一台游戏 PC 就能满足的配置。

另外:这是一个专用引擎,只跑这一个模型。如果你主要需求就是本地跑 Qwen3.8-Flash-Next,Strata 是目前最优化的方案;如果需要灵活支持多个模型,它不是正确的工具。

5500 stars 对于一个高度专用的 C++ 推理引擎来说是很强的社区反应,说明这个需求真实存在,且当前没有更好的替代方案。


开源仅供学习,MIT 协议,商业使用无限制。


🇬🇧 English

Strata: Run a 125B Model on a 12 GB GPU with MoE Expert Hot Cache

Open source for learning only: All projects discussed are from public repositories.


The Core Insight

A 125B parameter model normally needs multiple A100s or large server-class GPUs. Qwen3.8-Flash-Next is a 125B MoE model — 125B total parameters, but each token only activates a small subset of expert layers.

Strata builds an inference engine around this property: cache frequently used experts on the GPU, read infrequently used ones from RAM on demand. Result: 12 GB VRAM + 64 GB system RAM, up to 93 tokens/second (Q2_0 quantization).

Repo: github.com/Niko1221/Strata
5500+ stars | MIT | C++ | Author: Niko Veit (Niko1221)


How It Works

MoE inference property: MoE models don’t use all parameters for each token. Strata exploits this with heterogeneous storage:

  • Hot experts on GPU: Frequently called expert layers stay in VRAM for lowest latency
  • Cold experts in RAM: Infrequently used experts live in system memory, loaded on demand
  • Adaptive eviction: Exponential decay algorithm maintains the cache — hotter experts are harder to evict
  • Expert profile persistence: Expert usage distribution saved every 10 minutes and on exit; next startup loads the profile to skip cold-start

Cold experts are computed on CPU using AVX-512/AVX2 in parallel with GPU computation on hot experts — not a pure serial wait.

MTP speculative decoding: A small draft layer guesses the next several tokens; the main model verifies in batch. Correct guesses are accepted immediately; wrong ones fall back to normal generation. Delivers 1.6–1.8x additional speedup for short conversations (setting: MTP n-max 4).

Custom CUDA kernels: Rewritten specifically for Qwen3.8-Flash-Next’s architecture. Official YouTube titles claim “6x faster than llama.cpp” — take this with a grain of salt, as llama.cpp isn’t optimized for this style of expert offloading. But custom kernels do meaningfully outperform generic implementations.


Performance Numbers

Benchmarks from the author on RTX 5070 (12GB):

QuantizationShort context128K long context
Q2_0~93 t/s~74 t/s
IQ3_S~53 t/s~46 t/s

User-reported 75-96 t/s on RTX 5070 and RTX 3090 (24GB), consistent with official numbers.

Model file: ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF, ~75.8 GB total, stored in system memory and on disk — VRAM only holds the currently active expert cache.

Known performance regression: Issue #528 (v0.1.36, still under investigation) reports a 4-6x speed drop on continued conversations on RTX 5090. Short conversations are unaffected.


Installation

One-click install handles CUDA setup automatically:

# Linux
curl -fsSL https://raw.githubusercontent.com/Niko1221/Strata/main/install.sh | bash
# Windows: download and run install.bat

After startup:

  • Browser UI: Local chat interface
  • OpenAI-compatible API: http://localhost:PORT/v1/chat/completions
  • Anthropic-compatible API: http://localhost:PORT/v1/messages
  • Optional image input: Multimodal (note: mmproj adds ~907 MB VRAM overhead, reduces expert cache slots by 300-600 on 8-12 GB cards)

Where the Hardware Bar Actually Is

“12 GB GPU” is accurate, but the full requirements are:

HardwareRequirement
GPUNVIDIA RTX 30/40/50 series, ≥12 GB VRAM (8 GB fork exists)
System RAM64 GB — this is the harder bar to clear
Storage~80 GB free (model + engine)
OSWindows or Linux
BackendNVIDIA only (no AMD, no Mac, no CPU-only)

Most consumer gaming PCs ship with 16-32 GB RAM. 64 GB requires a deliberate upgrade. The GPU’s 12 GB VRAM is actually the easier constraint to satisfy.


Hard Constraints

Only runs Qwen3.8-Flash-Next. Strata’s CUDA kernels and caching logic are purpose-built for this one model architecture. Not a general-purpose inference engine.

NVIDIA-only. No AMD ROCm, no Apple Metal, no CPU-only mode.

Continued-conversation performance regression known (Issue #528, v0.1.36 under investigation). Short conversations unaffected.

MIT license — commercial use unrestricted.


Key Numbers

MetricValue
Stars5,500+
LicenseMIT
LanguageC++
Target modelQwen3.8-Flash-Next (125B MoE)
Min VRAM12 GB NVIDIA
Min system RAM64 GB
Model size~75.8 GB
Peak speed (Q2_0)~93 t/s (short ctx, RTX 5070)
Current versionv0.1.36
PlatformsWindows + Linux

Verdict

What Strata is doing is technically sound: MoE models are naturally suited to heterogeneous storage inference — each token activates only a fraction of experts, so you don’t need all weights on GPU. Strata engineers this into practice with custom CUDA kernels and speculative decoding, pushing a 125B model into consumer hardware territory.

The speed claims are real. 93 t/s for a 125B model is genuinely fast. But “12 GB GPU” omits an equally important condition: “plus 64 GB system RAM.” That’s not a standard gaming PC configuration — it’s a deliberate hardware investment.

Additional context: this is a single-model specialized engine. If your primary need is local Qwen3.8-Flash-Next inference, Strata is currently the most optimized option. If you need flexible multi-model support, this isn’t the right tool.

5,500+ stars for a highly specialized C++ inference engine is strong community signal: the demand is real, and there’s currently no better alternative for this specific use case.


Open source for learning only. MIT license — commercial use unrestricted.

💬 评论与讨论

使用 GitHub 账号登录后发表评论

关于本站 · 免责声明

🍄 Mushroom Research Blog 是非营利、免费公开的个人科技观察博客与公众号 XStack18,不接受商业合作、不代表任何企业或机构立场,也不谋求商业利益。我们以个人视角客观中立地记录和分析 AI、Web3 等领域的最新模型发布与技术动态——不止转述新闻标题或二手信息,而是给出有独立思考的深入分析,希望帮更多人获得有价值的一手科技认知。

⚠️ 文中介绍的开源代码与模型,仅供学习交流与技术借鉴。它们大多仍处于早期阶段,有待进一步研究和验证,请勿直接用于工作或生产环境;如需采用,请先自行充分测试,并核实其许可证与安全性。
Open-source code and models featured here are shared for learning and reference only. Most are early-stage and still need further study and verification — please don't use them directly in your work or in production. Test them thoroughly and check their licenses and security first.

  1. 本站文章均为作者基于公开信息的个人研究与观点整理,不代表文中提及的任何公司、产品、模型的官方立场,未与其构成商业关联或合作关系。
  2. 科技行业信息更新极快,我们尽力保证内容准确、及时,但不对完整性、实时性做绝对保证,具体请以相关企业/项目官方公告为准。
  3. 文中引用的第三方商标、产品名称、图片、数据等版权归原权利人所有,我们会尽量注明来源;如你认为存在版权疑问或侵权,请通过下方邮箱联系我们,收到通知后会尽快核实处理(更正、加注来源或删除)。
  4. 文章内容仅为技术科普与个人观点,不构成投资、法律或其他专业建议,据此进行任何决策的后果需自行判断和承担。

📮 侵权 / 勘误 / 合作咨询:[email protected]