HuggingFace 近三个月趋势:大模型抢头条,小模型统治实际使用

HuggingFace Q3 2026 Trending: Big Models Make Headlines, Small Models Run the World

Tech-News #HuggingFace#趋势分析#小模型#GGUF#本地推理#Qwen#开源生态
🇨🇳 中文

数据来源:HuggingFace 夏季 2026 开放模型状态报告、agents-radar 日报归档、Tech AI Magazine 月度整理。


两件事同时在发生

过去三个月,HuggingFace 上有两条完全不同的叙事在同步推进。

上层叙事:前沿大模型持续刷新规模上限。GPT-OSS 120B(7月)、Gemma 4 多模态(8月)、GLM-5.2 MoE + Kimi-K3 长文档(9月)、DeepSeek-V4-Flash 低延迟推理——每个月都有”新高”。

下层叙事:实际被下载和使用的,绝大多数是你从没听说过名字的小模型。

HuggingFace 自己在夏季报告里给出了数字:小于 1B 参数的模型,占 Hub 全时段总下载量的 83%。大于 100B 的,只有 1%。

这不是说大模型没人用——而是说,当我们谈”AI 社区在用什么”,我们主要在谈小模型。


三个月热榜速览

7 月:前沿模型集体登场

7 月最大的事是开源/开权重大模型集体更新:

模型参数特点
openai/gpt-oss-120b120B首次开放权重
deepseek-ai/DeepSeek-V3.2MoE更新版本
moonshotai/Kimi-K2-Instruct-长上下文
zai-org/GLM-4.6-中文优势
Qwen/Qwen3-Coder-480B-A35BMoE代码专精
black-forest-labs/FLUX.1-Kontext-dev-图像编辑

同月,openai/whisper-large-v3-turbo(ASR)和 nomic-ai/nomic-embed-text-v2-moe(Embedding)也在榜。工具类模型和生成类模型并驾齐驱。

8 月:多模态与速度

Gemma 4 是 8 月的主角——Google 的全模态模型,覆盖文字、图像、视频和原生音频,从轻量边缘版到企业 MoE 版都有。DeepSeek-V4-Flash 主打低延迟推理,在开源社区里流行度持续上升。

9 月:MoE + 长文档

GLM-5.2 以混合专家架构重新进入热榜视野,Kimi-K3 以长文档处理能力(法律合同、学术论文、知识管理)获得关注。两者都对应具体业务场景,不只是 benchmark 数字。

10 月初(今日):三个黑马

今天 agents-radar 日报里,前三名不是大模型:

Contrastive-LM/CLM-v0.1-8B(趋势分 485)
8B 参数的对比学习文本排序与验证模型。定位是 RAG pipeline 里的 reranker 和内容过滤层,不是生成模型。626 赞,2720 下载,属于”看起来没故事,但实际有用途”的那类模型。

TaichuAI/ZDTaichu5.0-9B(趋势分 475)
9B 视觉语言模型,专精空间推理——机器人、3D 场景理解、空间关系判断。2571 赞,12194 下载。中文团队出品,榜单上的稳定位。

nvidia/Nemotron-3-Diarization(趋势分 458)
100M 参数,专做说话人分离(最多 8 人)。M4 Mac 离线处理 10 分钟录音只需 3.1 秒。会议纪要、播客转写、客服录音标注的必备工具。26.4K 下载。

三个完全不同的细分任务,都是 10B 以下的专用模型。这是当日热榜给出的信号:开发者在为具体问题找具体工具,而不是找通用大模型。


小模型生态的真实规模

HuggingFace 夏季报告的另一组数字值得单独说:

GGUF 仓库增长 464%

相比之下,核心 transformers 库的仓库增长是 16%。GGUF 是专为本地推理设计的量化格式——这个增速意味着,“让模型在本地能跑”的需求,正在以压倒性的速度增长,远超新模型的发布速度本身。

unsloth/Qwen3.8-27B-GGUF 一个仓库就有超过 1100 万次下载。这不是在下载”开源大模型”,这是在下载”在我的机器上能实际运行的量化版本”。

Qwen 的社区主导

Qwen 系列在 Hub 上衍生出超过 15.1 万个模型,是 Meta(Llama 系列)衍生数量的 2.6 倍。Qwen GGUF 格式每月下载量 3960 万次,是 Llama GGUF 的 5 倍以上。

这意味着 Qwen 已经不只是一个模型系列,而是当前开源社区微调和量化的默认底座。当开发者想做一个垂直领域的小模型,第一个想到的基座大概率是 Qwen。


两套逻辑,同时成立

理解当前开源模型生态,需要接受两套逻辑同时成立的现实:

发布侧的逻辑:参数量和 benchmark 分数是竞争维度。大模型公司需要用每次新发布证明自己仍在前沿,GPT-OSS、Gemma 4、Kimi-K3 都是这个逻辑下的产物。这是新闻,也是必要的技术进步。

使用侧的逻辑:能不能在我的机器上跑,是不是能解决我的具体问题,模型有多小、推理有多快。83% 的下载量在 1B 以下,464% 的 GGUF 仓库增长,都在说同一件事:大多数开发者在意的是可用性,不是能力上限。

今日热榜的 CLM-v0.1-8B 是一个典型——没有参数大战,没有 benchmark 刷新,只是一个在 RAG pipeline 里做 reranking 的 8B 专用模型,却在当天排在趋势榜前三。这种模型不会上技术媒体的头版,但它会被几千个实际在做产品的开发者安静地下载和部署。


要关注的几个细分方向

根据三个月的热榜,以下几类模型反复出现:

RAG 工具链:reranker、embedding 模型持续上榜,说明 RAG 应用规模在扩大,工具链需求被具体化了。

音视频处理:Nemotron-3 说话人分离、Whisper turbo、腾讯 AuK 语音——实时/离线音频处理需求形成了稳定的细分赛道。

空间和多模态推理:ZDTaichu5.0-9B 的上榜是一个信号,空间理解(不只是图文问答)开始成为独立的能力评估维度,背后可能是机器人/工业自动化的需求在驱动。

Qwen 衍生微调:量化和微调的 Qwen 变体,包括各种语言、各种垂直领域,是社区活跃度的主要贡献来源。


关键数字汇总

指标数据
<1B 模型占总下载量83%
>100B 模型占总下载量1%
GGUF 仓库增长(年化)+464%
transformers 仓库增长+16%
Qwen 衍生模型数量151,448
Qwen vs Meta 衍生比2.6x
Qwen GGUF 月下载3960万次
Qwen vs Llama GGUF 比5x
HuggingFace 数据集总量100万+
85.6% 模型的下载量<200次

数据来源:HuggingFace 夏季 2026 开放模型状态报告(huggingface.co/blog/state-of-open-models-summer-2026)、agents-radar 日报、Tech AI Magazine 月度整理。以上仅供技术趋势参考,不构成模型选型建议。


🇬🇧 English

Data sources: HuggingFace Summer 2026 State of Open Models report, agents-radar daily digests, Tech AI Magazine monthly rankings.


Two Stories Running in Parallel

The past three months on HuggingFace have been driven by two completely separate narratives.

The headline narrative: Frontier models keep pushing scale ceilings. GPT-OSS 120B (July), Gemma 4 multimodal (August), GLM-5.2 MoE + Kimi-K3 long-context (September), DeepSeek-V4-Flash — each month had a new “record.”

The usage narrative: What’s actually being downloaded and run is almost entirely models you’ve never heard of — and most of them are small.

HuggingFace’s Summer 2026 State of Open Models report gives the number directly: models under 1B parameters account for 83% of all-time downloads on the Hub. Models over 100B account for 1%.

This isn’t a story about big models being irrelevant. It’s a story about what “the AI community” actually uses every day.


July: Frontier model wave

GPT-OSS 120B (first open-weights from OpenAI), DeepSeek-V3.2, Kimi-K2-Instruct, GLM-4.6, Qwen3-Coder-480B-A35B (code-specialized MoE), FLUX.1-Kontext-dev (image editing), whisper-large-v3-turbo, nomic-embed-text-v2-moe. Tool-class models trending alongside generative ones.

August: Multimodal and speed

Gemma 4 dominated — Google’s full-modality model handling text, image, video, and native audio, from lightweight edge variants to enterprise MoE. DeepSeek-V4-Flash gained traction for low-latency inference.

September: MoE and long documents

GLM-5.2 re-entered attention via MoE architecture; Kimi-K3 for long-context enterprise use (legal contracts, academic literature, knowledge management). Both tied to concrete business scenarios, not just benchmarks.

Early October (today): Three surprise entries

Today’s top three on agents-radar aren’t large models:

Contrastive-LM/CLM-v0.1-8B (trend score 485): 8B contrastive learning model for text ranking and verification. Built for RAG pipelines as a reranker and content filter — not a generative model. 626 likes, 2720 downloads. The “no narrative but actually useful” category.

TaichuAI/ZDTaichu5.0-9B (trend score 475): 9B VLM specialized for spatial reasoning — robotics, 3D scene understanding, spatial relationship judgments. 2571 likes, 12194 downloads. Chinese team, consistent trending position.

nvidia/Nemotron-3-Diarization (trend score 458): 100M parameters for speaker diarization (up to 8 speakers). M4 Mac processes 10 minutes of audio offline in 3.1 seconds. 26.4K downloads.

Three completely different specialized tasks, all under 10B parameters. The signal: developers are looking for specific tools for specific problems, not general-purpose large models.


The Real Scale of the Small Model Ecosystem

GGUF repository growth: +464%

Core transformers library repos grew 16% over the same period. GGUF is the quantization format designed for local inference. This growth rate means “making models actually run locally” is growing at an overwhelming pace — far outpacing even the release of new models.

A single repo — unsloth/Qwen3.8-27B-GGUF — has over 11 million downloads. Developers aren’t downloading “a large open-source model.” They’re downloading “the quantized version that will actually run on my hardware.”

Qwen’s community dominance

Qwen has over 151,000 derivative models on the Hub — 2.6x Meta’s (Llama) total footprint. Qwen GGUF downloads reach 39.6 million per month — 5x Llama GGUF.

Qwen has become the default fine-tuning and quantization base for the open-source community. When a developer wants to build a vertical-domain small model, their first instinct is now Qwen.


Two Logics, Both True

Understanding the current open model ecosystem requires accepting that two things are simultaneously true:

The release side’s logic: Parameter count and benchmark scores are the competitive dimension. Big model labs need each release to demonstrate continued frontier relevance. GPT-OSS, Gemma 4, Kimi-K3 all live in this logic. This is news, and it represents genuine technical progress.

The usage side’s logic: Can it run on my machine? Does it solve my specific problem? How small is it, how fast does it run? 83% of downloads under 1B, 464% GGUF growth — both point to the same conclusion: most developers care about deployability, not capability ceiling.

CLM-v0.1-8B today is a case in point — no parameter competition, no benchmark record, just an 8B specialized reranker for RAG pipelines, quietly ranking top-3 on the trending digest. This kind of model won’t make tech media front pages. But it will be silently downloaded and deployed by thousands of developers actually building products.


Segments to Watch

Three months of trending data points to consistent signals in specific categories:

RAG toolchain: Rerankers and embedding models keep appearing. RAG applications are scaling up, and toolchain demand is becoming concrete and specialized.

Audio/video processing: Nemotron-3 diarization, Whisper turbo, TencentAuK speech — real-time and offline audio processing has become a stable niche with consistent demand.

Spatial and multimodal reasoning: ZDTaichu5.0-9B’s appearance is a signal that spatial understanding (not just image-text Q&A) is becoming an independently evaluated capability — likely driven by robotics and industrial automation demand.

Qwen derivative fine-tunes: Quantized and fine-tuned Qwen variants across languages and vertical domains are the primary contributor to community activity volume.


Key Numbers

MetricValue
<1B models share of total downloads83%
>100B models share1%
GGUF repository growth+464%
Transformers library growth+16%
Qwen derivative models151,448
Qwen vs Meta footprint2.6x
Qwen GGUF monthly downloads39.6M
Qwen vs Llama GGUF5x
HuggingFace total datasets1M+
Models with <200 downloads85.6%

Data from the HuggingFace Summer 2026 State of Open Models report, agents-radar daily digests, and Tech AI Magazine monthly rankings. For technical trend reference only.

💬 评论与讨论

使用 GitHub 账号登录后发表评论

关于本站 · 免责声明

🍄 Mushroom Research Blog 是非营利、免费公开的个人科技观察博客与公众号 XStack18,不接受商业合作、不代表任何企业或机构立场,也不谋求商业利益。我们以个人视角客观中立地记录和分析 AI、Web3 等领域的最新模型发布与技术动态——不止转述新闻标题或二手信息,而是给出有独立思考的深入分析,希望帮更多人获得有价值的一手科技认知。

⚠️ 文中介绍的开源代码与模型,仅供学习交流与技术借鉴。它们大多仍处于早期阶段,有待进一步研究和验证,请勿直接用于工作或生产环境;如需采用,请先自行充分测试,并核实其许可证与安全性。
Open-source code and models featured here are shared for learning and reference only. Most are early-stage and still need further study and verification — please don't use them directly in your work or in production. Test them thoroughly and check their licenses and security first.

  1. 本站文章均为作者基于公开信息的个人研究与观点整理,不代表文中提及的任何公司、产品、模型的官方立场,未与其构成商业关联或合作关系。
  2. 科技行业信息更新极快,我们尽力保证内容准确、及时,但不对完整性、实时性做绝对保证,具体请以相关企业/项目官方公告为准。
  3. 文中引用的第三方商标、产品名称、图片、数据等版权归原权利人所有,我们会尽量注明来源;如你认为存在版权疑问或侵权,请通过下方邮箱联系我们,收到通知后会尽快核实处理(更正、加注来源或删除)。
  4. 文章内容仅为技术科普与个人观点,不构成投资、法律或其他专业建议,据此进行任何决策的后果需自行判断和承担。

📮 侵权 / 勘误 / 合作咨询:[email protected]