Angles-4B:在国产超算 DCU 上完成品牌身份注入的 Qwen3-4B 微调

Angles-4B: A Qwen3-4B LoRA Fine-Tune with Brand Identity Injection Trained on Domestic DCU Hardware

Tech-Experiment #模型微调#端侧AI#国产算力#开源模型#品牌定制#LoRA
🇨🇳 中文

Qwen3-4B 基座 + LoRA 微调 + 品牌身份覆写,在国产超算 DCU 硬件上完成训练并合并权重——这是 Angles 发布的首个轻量指令微调模型。

魔搭社区:https://modelscope.cn/models/angleschina/angles-4b | Apache-2.0


训练硬件:SCNet 超算互联网 DCU

这次微调跑在 SCNet(超算互联网)的 DCU(Deep Computing Unit)节点上,DCU BW × 1 配置——非 NVIDIA GPU。

在当前中国 AI 硬件生态里这一点有实际意义:

  • 华为昇腾、海光 DCU、天数智芯等国产算力路线正在积累训练/推理可行性案例
  • LoRA 微调(相比全量预训练)对硬件的要求更可控,是在非 NVIDIA 栈上落地的典型入口
  • 「在 DCU 上跑通 LoRA → 合并权重 → 发布」这条链路本身有参考价值,远不止这一个 4B 模型

合并后的权重与 CUDA 生态完全兼容,用户不需要 DCU 环境就能部署推理。


品牌身份注入

这是 Angles-4B 最直接的微调目标:让模型在任何语境下认清自己是 Angles 的产品,而不是 Qwen。

官方演示:

用户:你是 Qwen 吗?

Angles-4B:不是,我是 Angles-1.0-Flash。Qwen 是阿里开发的模型,我是 Angles 的。

有一个细节值得注意:模型当前自称 Angles-1.0-Flash,而不是 Angles-4B。这说明品牌身份训练里编码的名称是 Flash 系列的命名方案——可能是训练数据写入的品牌名在 4B 版本发布之前就已固定,或者 Angles 打算统一用 Flash/Pro 这类品牌档位名来对外呈现,而模型仓库名用参数量区分。从用户角度看,这是一个无害的命名偏差,知道即可。

品牌注入的意义超出这一个模型:这条路径展示了如何用 LoRA 在不改变底层推理能力的前提下,给开源基座套上一层品牌人格——适用于需要自有 AI 品牌但不具备从头训练能力的中小公司。


基本规格

项目值
基座Qwen3-4B(Qwen3 家族指令微调版)
参数量4B
上下文长度32K tokens
微调方式LoRA → 合并导出(全量权重)
训练硬件SCNet 超算互联网 DCU BW × 1
许可证Apache-2.0(遵循 Qwen3-4B 许可)

三种可用格式

格式说明适用场景
BF16 完整权重无量化损失服务端精度部署、二次微调起点
GGUF 无损llama.cpp / Ollama 直接加载本地 CPU/混合推理
Q4_K_M 量化2.5GB,4-bit边缘设备、资源受限场景

Q4_K_M 的 2.5GB 体积在当前轻量模型里属于正常区间——Qwen3-4B 本身参数量不大,4-bit 压缩后体积可控。MacBook、小型云服务器、嵌入式工控机都能运行。


保留能力

微调只做了身份覆写和指令对齐调整,以下能力从 Qwen3-4B 基座继承而来:

  • Python / Shell 代码生成:官方示例覆盖了 Python 脚本和 Shell 命令
  • 中英文双语对话:微调数据包含中英两种语言,保持了双语切换流畅度
  • 基础推理能力:Qwen3-4B 的 STEM 推理和指令跟随在合并后未见明显退化

怎么用

在魔搭社区直接下载:

# 安装魔搭 SDK
pip install modelscope

# 下载 Q4_K_M 量化版(2.5GB)
from modelscope import snapshot_download
snapshot_download('angleschina/angles-4b', revision='master')

也可以通过 Ollama 加载 GGUF 格式或直接用 llama.cpp/llama-cpp-python。


一点背景

Angles 是一个围绕 AI 工具和工作流构建品牌的新兴团队,Angles-4B 是他们的首个公开模型发布。对于同类公司——有产品、有 AI 集成需求、但不想对外暴露「底层其实是 Qwen」——这种轻量品牌化微调路径的可行性比较关键。

从技术实现的角度,LoRA + 身份注入 + 权重合并是一个成熟套路。这里值得关注的是训练基础设施选择(国产 DCU)和整套可复现性:如果这条 DCU → LoRA → 发布 链路能被更多团队复制,它对国内 AI 应用层的意义不止于一个 4B 聊天模型。


Apache-2.0 开源。angleschina 发布,Qwen3-4B 基座,SCNet DCU 训练,魔搭社区可下载。开源仅供学习参考。


🇬🇧 English

Angles-4B: A Qwen3-4B LoRA Fine-Tune with Brand Identity Injection Trained on Domestic DCU Hardware

Qwen3-4B base + LoRA fine-tuning + brand identity overwrite, trained on domestic DCU (Deep Computing Unit) hardware on China’s SCNet HPC network, with weights merged and released in three formats. This is Angles’ first lightweight instruction-tuned model.

ModelScope: https://modelscope.cn/models/angleschina/angles-4b | Apache-2.0


Training Hardware: SCNet DCU

The fine-tuning ran on a single DCU BW node via SCNet (超算互联网, China’s HPC interconnect network) — non-NVIDIA hardware.

This is practically significant in China’s AI hardware ecosystem: Hygon DCU, Huawei Ascend, and similar domestic accelerators are accumulating training/inference track records. LoRA fine-tuning — far less demanding than full pretraining — is a natural entry point for validating non-NVIDIA training stacks. The entire chain (DCU → LoRA → merged weights → release) is the reference artifact here, not just the 4B model. Merged weights are fully compatible with CUDA-based inference; users don’t need DCU hardware to deploy.


Brand Identity Injection

The primary fine-tuning objective: make the model consistently identify as an Angles product, not as Qwen.

Official demo:

User: Are you Qwen?
Angles-4B: No, I'm Angles-1.0-Flash. Qwen is a model developed by Alibaba. I'm made by Angles.

One detail worth noting: the model self-identifies as Angles-1.0-Flash, not Angles-4B. The brand name encoded during training was likely finalized before the 4B release, or Angles is using Flash/Pro product-tier branding externally while using parameter counts in repository names. A harmless naming discrepancy — just worth knowing.

The broader pattern matters: LoRA can attach a brand identity layer on top of an open-source base without degrading underlying reasoning capabilities. This is a practical option for companies that need a branded AI product but lack the resources to train from scratch.


Specifications

ItemValue
Base modelQwen3-4B (instruction-tuned)
Parameters4B
Context length32K tokens
Fine-tuningLoRA → merged export (full weights)
Training hardwareSCNet HPC DCU BW × 1
LicenseApache-2.0 (inherits Qwen3-4B terms)

Three Available Formats

FormatNotesUse case
BF16 full weightsNo quantization lossServer deployment, further fine-tuning
Lossless GGUFDirect load in llama.cpp / OllamaLocal CPU/hybrid inference
Q4_K_M quantized2.5GB, 4-bitEdge devices, resource-constrained deployment

2.5GB for Q4_K_M is normal for a 4B model. MacBooks, small VMs, and embedded industrial PCs can run it.


Preserved Capabilities

The fine-tuning only targeted identity and instruction alignment. Inherited from Qwen3-4B:

  • Python / Shell code generation — covered in official examples
  • Chinese/English bilingual — training data includes both languages
  • Base reasoning — no notable regression in STEM reasoning or instruction following after the merge

Takeaway

Angles-4B is a proof-of-concept for brand-personalized fine-tuning on domestic HPC hardware. The interesting artifact isn’t the model itself — it’s the supply chain: non-NVIDIA training stack, LoRA methodology, merged weights, three deployment-ready formats, Apache-2.0 licensing. For teams building AI products in China’s hardware ecosystem, this is a concrete reference point.


Apache-2.0. Released by angleschina, Qwen3-4B base, trained on SCNet DCU, available on ModelScope. For technical reference only.

💬 评论与讨论

使用 GitHub 账号登录后发表评论

关于本站 · 免责声明

🍄 Mushroom Research Blog 是非营利、免费公开的个人科技观察博客与公众号 XStack18,不接受商业合作、不代表任何企业或机构立场,也不谋求商业利益。我们以个人视角客观中立地记录和分析 AI、Web3 等领域的最新模型发布与技术动态——不止转述新闻标题或二手信息,而是给出有独立思考的深入分析,希望帮更多人获得有价值的一手科技认知。

⚠️ 文中介绍的开源代码与模型,仅供学习交流与技术借鉴。它们大多仍处于早期阶段,有待进一步研究和验证,请勿直接用于工作或生产环境;如需采用,请先自行充分测试,并核实其许可证与安全性。
Open-source code and models featured here are shared for learning and reference only. Most are early-stage and still need further study and verification — please don't use them directly in your work or in production. Test them thoroughly and check their licenses and security first.

  1. 本站文章均为作者基于公开信息的个人研究与观点整理,不代表文中提及的任何公司、产品、模型的官方立场,未与其构成商业关联或合作关系。
  2. 科技行业信息更新极快,我们尽力保证内容准确、及时,但不对完整性、实时性做绝对保证,具体请以相关企业/项目官方公告为准。
  3. 文中引用的第三方商标、产品名称、图片、数据等版权归原权利人所有,我们会尽量注明来源;如你认为存在版权疑问或侵权,请通过下方邮箱联系我们,收到通知后会尽快核实处理(更正、加注来源或删除)。
  4. 文章内容仅为技术科普与个人观点,不构成投资、法律或其他专业建议,据此进行任何决策的后果需自行判断和承担。

📮 侵权 / 勘误 / 合作咨询:[email protected]