Kandinsky 6.0 Video:双流 CrossDiT,音视频同步生成开源了
Kandinsky 6.0 Video: Dual-Stream CrossDiT for Synchronized Audio-Video Generation
一句话同时出画面和声音,不是先渲视频再配音。
Kandinsky Lab(隶属 Sber)在 2026 年 10 月 4 日发布 Kandinsky 6.0 Video,底层架构叫 CrossDiT:视频流和音频流同时推理,通过双向交叉注意力在每个时间步交换信息,最终输出的画面和声音是共同生长出来的,而不是拼接出来的。
arXiv: https://arxiv.org/abs/2610.05608 | MIT | Kandinsky Lab (Sber)
两种生成模式
T2AV(文生音视频):一句文字提示,直接出 5 秒带声音的视频。提示里写的内容——场景、动作、音效——会同时在画面和音轨里体现。
I2AV(图生音视频):给一张图作为第一帧,模型推测接下来会发生什么,自动生成画面运动和配套音频。适合从静态图或参考概念图出发做创作。
不需要声音时可以单独关闭音频输出,只要画面。
CrossDiT 架构
传统音视频生成的做法是把两个独立模型串联:先视频,再配音。CrossDiT 做的是并行双流:
- 视频流:从预训练的 Kandinsky 5.0 视频基础模型初始化,不从零训练
- 音频流:全新从零构建,先在大规模音频语料上预训练,再与视频流联合优化
- 双向交叉注意力:视频流的每一层可以”看到”音频流当前状态,反之亦然——帧和声音互相影响,保持同步
这是训练设计上的选择,不是推理时做的后期对齐。
规格
| 项目 | 参数 |
|---|---|
| 视频时长 | 5 秒(121 帧) |
| 帧率 | 24 fps |
| 音频采样率 | 44 kHz |
| 唇形同步 | ✅ 原生支持 |
| 输出分辨率 | 训练分辨率(+VSR 可升 Full HD) |
每个尺寸都有三种变体:
- base:基础检查点
- pretrained:更长训练周期版本
- distilled:10 步蒸馏版,推理更快
两个尺寸
Lite(3B):面向本地部署和快速迭代,计算要求低。
Pro(29B):面向质量优先场景,Sber 内部人工评测结果是:在语音质量、唇形同步准确率、音视频对齐度上压过了 Kling 2.6、Veo 3.1 Fast、MiniMax H3、Seedance 2.0。评测基准是 VABench(专门针对音视频联合生成的评测集)。
⚠️ 这是 Sber 自己发布的对比数据,未经独立第三方复现。评测框架 VABench 也是 Kandinsky Lab 配套发布的,方法论中立性待社区验证。
训练流程
音频流预训练(大规模音频语料)
↓
视频+音频联合训练(配对音视频数据)
↓
有监督微调(SFT)
↓
强化学习后训练(RL)
↓
蒸馏(10步快速版本)
音频流不依赖任何现有的开源音频模型——这是完全自研的新流,专门针对音视频联合生成优化。
超分辨率模块:VSR
Kandinsky 6.0 VSR(1.4B 独立模型):将生成的视频升分辨率到 Full HD(1920×1080)。可以在 CrossDiT 主生成结束后单独调用,也可以集成进端到端流水线。
部署和体验
免费体验:通过 Sber 的 GigaChat 服务,不需要部署,直接在线试用。
本地部署:
- GitHub(参考):https://github.com/kandinskylab/kandinsky-6
- HuggingFace 检查点:
kandinskylab/Kandinsky-6.0-Pro-5s-Diffusers - ComfyUI 集成已有社区支持
一句话说清楚
Kandinsky 6.0 Video 不是音频条件视频生成,也不是视频条件音频生成——是两者联合生成,CrossDiT 在扩散过程里让音频流和视频流互相影响,最终输出的声画是共同演化出来的。MIT 开源,Pro 29B 在 VABench 上排首位(供应商自测)。本地可用 Lite 3B,追求质量走 Pro,免部署直接去 GigaChat。
MIT 许可。Kandinsky Lab(Sber)2026-10-04 发布,arXiv 2610.05608。开源仅供学习参考。
Kandinsky 6.0 Video: Dual-Stream CrossDiT for Synchronized Audio-Video Generation
One text prompt. Video and audio at the same time — not sequenced, but jointly generated.
Kandinsky Lab (Sber) released Kandinsky 6.0 Video on October 4, 2026. The architecture, called CrossDiT, runs a video stream and an audio stream in parallel, connected at every timestep via bidirectional cross-attention. The audio and video literally evolve together during diffusion — they’re not stitched together in post-processing.
arXiv: https://arxiv.org/abs/2610.05608 | MIT | Kandinsky Lab (Sber)
Two Input Modes
T2AV (Text-to-Audio-Video): A text prompt generates a 5-second clip with synchronized audio. Whatever the prompt describes — the scene, the motion, the ambient sound — appears in both the visual track and the audio track.
I2AV (Image-to-Audio-Video): A reference image becomes the first frame. The model infers what happens next and generates both the motion and the audio track. Good for starting from a concept image or storyboard.
Audio can be disabled to output video-only.
CrossDiT Architecture
Most audio-video generation pipelines chain two independent models: generate video, then dub audio. CrossDiT uses parallel dual streams:
- Video stream: initialized from the pretrained Kandinsky 5.0 video foundation model
- Audio stream: built from scratch, first pretrained on large audio corpora, then jointly optimized with the video stream
- Bidirectional cross-attention: every layer in the video stream can “see” the audio stream state, and vice versa — frames and sound influence each other, maintaining sync
This sync is a training design choice, not post-hoc alignment.
Specs
| Item | Value |
|---|---|
| Clip length | 5 seconds (121 frames) |
| Frame rate | 24 fps |
| Audio sample rate | 44 kHz |
| Lip sync | ✅ Native |
| Base resolution | Training resolution (VSR optional for Full HD) |
Three variants per size:
- base: standard checkpoint
- pretrained: extended training
- distilled: 10-step fast inference
Two Sizes
Lite (3B): For local deployment and fast iteration.
Pro (29B): Quality-first. Per Sber’s human evaluation on VABench, Pro ranks first on speech quality, lip-sync accuracy, and audio-video alignment — ahead of Kling 2.6, Veo 3.1 Fast, MiniMax H3, and Seedance 2.0.
⚠️ This comparison data comes from the model’s authors. VABench was also released by Kandinsky Lab, so the methodology hasn’t been independently verified yet.
Training Pipeline
Audio stream pretrain (large-scale audio corpora)
↓
Joint video+audio training (paired AV data)
↓
Supervised fine-tuning (SFT)
↓
RL post-training
↓
Distillation (10-step fast variant)
The audio stream doesn’t depend on any existing open-source audio model — it’s purpose-built for joint AV generation.
VSR Super-Resolution Module
Kandinsky 6.0 VSR (1.4B standalone model): upscales generated video to Full HD (1920×1080). Can run after CrossDiT or be integrated into an end-to-end pipeline.
TL;DR
Kandinsky 6.0 Video isn’t audio-conditioned video generation or video-conditioned audio generation — it’s joint generation: CrossDiT runs both streams in parallel through the diffusion process, cross-attending at every step. MIT license. Pro 29B leads VABench (vendor self-test). Lite 3B for local use, GigaChat for zero-setup demo.
MIT license. Released by Kandinsky Lab (Sber) on 2026-10-04. arXiv 2610.05608. For technical reference only.
关于本站 · 免责声明
🍄 Mushroom Research Blog 是非营利、免费公开的个人科技观察博客与公众号 XStack18,不接受商业合作、不代表任何企业或机构立场,也不谋求商业利益。我们以个人视角客观中立地记录和分析 AI、Web3 等领域的最新模型发布与技术动态——不止转述新闻标题或二手信息,而是给出有独立思考的深入分析,希望帮更多人获得有价值的一手科技认知。
⚠️ 文中介绍的开源代码与模型,仅供学习交流与技术借鉴。它们大多仍处于早期阶段,有待进一步研究和验证,请勿直接用于工作或生产环境;如需采用,请先自行充分测试,并核实其许可证与安全性。
Open-source code and models featured here are shared for learning and reference only. Most are early-stage and still need further study and verification — please don't use them directly in your work or in production. Test them thoroughly and check their licenses and security first.
- 本站文章均为作者基于公开信息的个人研究与观点整理,不代表文中提及的任何公司、产品、模型的官方立场,未与其构成商业关联或合作关系。
- 科技行业信息更新极快,我们尽力保证内容准确、及时,但不对完整性、实时性做绝对保证,具体请以相关企业/项目官方公告为准。
- 文中引用的第三方商标、产品名称、图片、数据等版权归原权利人所有,我们会尽量注明来源;如你认为存在版权疑问或侵权,请通过下方邮箱联系我们,收到通知后会尽快核实处理(更正、加注来源或删除)。
- 文章内容仅为技术科普与个人观点,不构成投资、法律或其他专业建议,据此进行任何决策的后果需自行判断和承担。
📮 侵权 / 勘误 / 合作咨询:[email protected]
💬 评论与讨论
使用 GitHub 账号登录后发表评论