Kandinsky 6.0 Video:双流 CrossDiT,音视频同步生成开源了

Kandinsky 6.0 Video: Dual-Stream CrossDiT for Synchronized Audio-Video Generation

Tech-News #视频生成#音频生成#开源模型#多模态#AI工具
🇨🇳 中文

一句话同时出画面和声音,不是先渲视频再配音。

Kandinsky Lab(隶属 Sber)在 2026 年 10 月 4 日发布 Kandinsky 6.0 Video,底层架构叫 CrossDiT:视频流和音频流同时推理,通过双向交叉注意力在每个时间步交换信息,最终输出的画面和声音是共同生长出来的,而不是拼接出来的。

arXiv: https://arxiv.org/abs/2610.05608 | MIT | Kandinsky Lab (Sber)


两种生成模式

T2AV(文生音视频):一句文字提示,直接出 5 秒带声音的视频。提示里写的内容——场景、动作、音效——会同时在画面和音轨里体现。

I2AV(图生音视频):给一张图作为第一帧,模型推测接下来会发生什么,自动生成画面运动和配套音频。适合从静态图或参考概念图出发做创作。

不需要声音时可以单独关闭音频输出,只要画面。


CrossDiT 架构

传统音视频生成的做法是把两个独立模型串联:先视频,再配音。CrossDiT 做的是并行双流:

  • 视频流:从预训练的 Kandinsky 5.0 视频基础模型初始化,不从零训练
  • 音频流:全新从零构建,先在大规模音频语料上预训练,再与视频流联合优化
  • 双向交叉注意力:视频流的每一层可以”看到”音频流当前状态,反之亦然——帧和声音互相影响,保持同步

这是训练设计上的选择,不是推理时做的后期对齐。


规格

项目参数
视频时长5 秒(121 帧)
帧率24 fps
音频采样率44 kHz
唇形同步✅ 原生支持
输出分辨率训练分辨率(+VSR 可升 Full HD)

每个尺寸都有三种变体:

  • base:基础检查点
  • pretrained:更长训练周期版本
  • distilled:10 步蒸馏版,推理更快

两个尺寸

Lite(3B):面向本地部署和快速迭代,计算要求低。

Pro(29B):面向质量优先场景,Sber 内部人工评测结果是:在语音质量、唇形同步准确率、音视频对齐度上压过了 Kling 2.6、Veo 3.1 Fast、MiniMax H3、Seedance 2.0。评测基准是 VABench(专门针对音视频联合生成的评测集)。

⚠️ 这是 Sber 自己发布的对比数据,未经独立第三方复现。评测框架 VABench 也是 Kandinsky Lab 配套发布的,方法论中立性待社区验证。


训练流程

音频流预训练(大规模音频语料)
        ↓
视频+音频联合训练(配对音视频数据)
        ↓
有监督微调(SFT)
        ↓
强化学习后训练(RL)
        ↓
蒸馏(10步快速版本)

音频流不依赖任何现有的开源音频模型——这是完全自研的新流,专门针对音视频联合生成优化。


超分辨率模块:VSR

Kandinsky 6.0 VSR(1.4B 独立模型):将生成的视频升分辨率到 Full HD(1920×1080)。可以在 CrossDiT 主生成结束后单独调用,也可以集成进端到端流水线。


部署和体验

免费体验:通过 Sber 的 GigaChat 服务,不需要部署,直接在线试用。

本地部署:


一句话说清楚

Kandinsky 6.0 Video 不是音频条件视频生成,也不是视频条件音频生成——是两者联合生成,CrossDiT 在扩散过程里让音频流和视频流互相影响,最终输出的声画是共同演化出来的。MIT 开源,Pro 29B 在 VABench 上排首位(供应商自测)。本地可用 Lite 3B,追求质量走 Pro,免部署直接去 GigaChat。


MIT 许可。Kandinsky Lab(Sber)2026-10-04 发布,arXiv 2610.05608。开源仅供学习参考。


🇬🇧 English

Kandinsky 6.0 Video: Dual-Stream CrossDiT for Synchronized Audio-Video Generation

One text prompt. Video and audio at the same time — not sequenced, but jointly generated.

Kandinsky Lab (Sber) released Kandinsky 6.0 Video on October 4, 2026. The architecture, called CrossDiT, runs a video stream and an audio stream in parallel, connected at every timestep via bidirectional cross-attention. The audio and video literally evolve together during diffusion — they’re not stitched together in post-processing.

arXiv: https://arxiv.org/abs/2610.05608 | MIT | Kandinsky Lab (Sber)


Two Input Modes

T2AV (Text-to-Audio-Video): A text prompt generates a 5-second clip with synchronized audio. Whatever the prompt describes — the scene, the motion, the ambient sound — appears in both the visual track and the audio track.

I2AV (Image-to-Audio-Video): A reference image becomes the first frame. The model infers what happens next and generates both the motion and the audio track. Good for starting from a concept image or storyboard.

Audio can be disabled to output video-only.


CrossDiT Architecture

Most audio-video generation pipelines chain two independent models: generate video, then dub audio. CrossDiT uses parallel dual streams:

  • Video stream: initialized from the pretrained Kandinsky 5.0 video foundation model
  • Audio stream: built from scratch, first pretrained on large audio corpora, then jointly optimized with the video stream
  • Bidirectional cross-attention: every layer in the video stream can “see” the audio stream state, and vice versa — frames and sound influence each other, maintaining sync

This sync is a training design choice, not post-hoc alignment.


Specs

ItemValue
Clip length5 seconds (121 frames)
Frame rate24 fps
Audio sample rate44 kHz
Lip sync✅ Native
Base resolutionTraining resolution (VSR optional for Full HD)

Three variants per size:

  • base: standard checkpoint
  • pretrained: extended training
  • distilled: 10-step fast inference

Two Sizes

Lite (3B): For local deployment and fast iteration.

Pro (29B): Quality-first. Per Sber’s human evaluation on VABench, Pro ranks first on speech quality, lip-sync accuracy, and audio-video alignment — ahead of Kling 2.6, Veo 3.1 Fast, MiniMax H3, and Seedance 2.0.

⚠️ This comparison data comes from the model’s authors. VABench was also released by Kandinsky Lab, so the methodology hasn’t been independently verified yet.


Training Pipeline

Audio stream pretrain (large-scale audio corpora)
        ↓
Joint video+audio training (paired AV data)
        ↓
Supervised fine-tuning (SFT)
        ↓
RL post-training
        ↓
Distillation (10-step fast variant)

The audio stream doesn’t depend on any existing open-source audio model — it’s purpose-built for joint AV generation.


VSR Super-Resolution Module

Kandinsky 6.0 VSR (1.4B standalone model): upscales generated video to Full HD (1920×1080). Can run after CrossDiT or be integrated into an end-to-end pipeline.


TL;DR

Kandinsky 6.0 Video isn’t audio-conditioned video generation or video-conditioned audio generation — it’s joint generation: CrossDiT runs both streams in parallel through the diffusion process, cross-attending at every step. MIT license. Pro 29B leads VABench (vendor self-test). Lite 3B for local use, GigaChat for zero-setup demo.


MIT license. Released by Kandinsky Lab (Sber) on 2026-10-04. arXiv 2610.05608. For technical reference only.

💬 评论与讨论

使用 GitHub 账号登录后发表评论

关于本站 · 免责声明

🍄 Mushroom Research Blog 是非营利、免费公开的个人科技观察博客与公众号 XStack18,不接受商业合作、不代表任何企业或机构立场,也不谋求商业利益。我们以个人视角客观中立地记录和分析 AI、Web3 等领域的最新模型发布与技术动态——不止转述新闻标题或二手信息,而是给出有独立思考的深入分析,希望帮更多人获得有价值的一手科技认知。

⚠️ 文中介绍的开源代码与模型,仅供学习交流与技术借鉴。它们大多仍处于早期阶段,有待进一步研究和验证,请勿直接用于工作或生产环境;如需采用,请先自行充分测试,并核实其许可证与安全性。
Open-source code and models featured here are shared for learning and reference only. Most are early-stage and still need further study and verification — please don't use them directly in your work or in production. Test them thoroughly and check their licenses and security first.

  1. 本站文章均为作者基于公开信息的个人研究与观点整理,不代表文中提及的任何公司、产品、模型的官方立场,未与其构成商业关联或合作关系。
  2. 科技行业信息更新极快,我们尽力保证内容准确、及时,但不对完整性、实时性做绝对保证,具体请以相关企业/项目官方公告为准。
  3. 文中引用的第三方商标、产品名称、图片、数据等版权归原权利人所有,我们会尽量注明来源;如你认为存在版权疑问或侵权,请通过下方邮箱联系我们,收到通知后会尽快核实处理(更正、加注来源或删除)。
  4. 文章内容仅为技术科普与个人观点,不构成投资、法律或其他专业建议,据此进行任何决策的后果需自行判断和承担。

📮 侵权 / 勘误 / 合作咨询:[email protected]