腾讯 Youtu-Parsing-Omni:一个 5B 模型解析文档、图表、音频、视频

Tencent Youtu-Parsing-Omni: One 5B Model to Parse Documents, Charts, Audio, and Video

Research #多模态#文档解析#OCR#腾讯#开源模型
🇨🇳 中文

大多数解析工具要么处理文档,要么处理图像,要么处理音频——分门别类,各司其职。

Youtu-Parsing-Omni 想把这件事做成一个:一个 5B 模型,一个统一 JSON,八种输入模态。

GitHub: https://github.com/TencentCloudADP/youtu-parsing HuggingFace: https://huggingface.co/tencent/Youtu-Parsing-Omni

88 stars,腾讯 Youtu Lab 出品,TencentCloudADP 发布。


八种模态,一个出口

--task 参数决定任务类型:

任务输入
document文档页(PDF、扫描件)
natural_image自然图像或照片
graphics_chart柱状图/折线图/饼图
graphics_flowchart流程图和示意图
graphics_geometric几何图(数学/科学题)
audio音频片段
natural_video自然视频(无大量覆盖文字)
textrich_video含文字视频(课件、录屏)

所有任务输出同一个 OmniSchema JSON,内部字段按任务类型填充,结构如下:

{
  "modality": "image",
  "subtype": "document",
  "elements": [
    {
      "type": "text",
      "content": "...",
      "bbox": "<box><x_85><y_275><x_1000><y_1000></box>"
    }
  ],
  "reading_order": [0, 1, 2],
  "global_description": "..."
}

音视频输出中 segments 字段携带 HH:MM:SS 时间戳。


感知层输出什么

输入类型感知输出
文档版面块 + OCR + bbox + 表格(OTSL) + 公式(LaTeX) + 阅读顺序
图表Markdown 表(图表)、Mermaid(流程图)、点线弧测量(几何)
音频时间戳 + 说话人 ASR + 非语音声学事件标签
自然视频时间戳 + 场景描述 + 镜头运动标签
含文字视频时间戳 + ASR + OCR + structured_report(Markdown 结构化摘要)

bbox 用特殊 token 表示,坐标系归一化到 0–1000 网格:<box><x_85><y_275><x_1000><y_1000></box>


认知层输出什么

每个任务的输出里都有 global_description,内容随任务类型变化:

  • 图表/文档 → 内容摘要
  • 自然图像 → 场景描述
  • 自然视频 → 跨时序的动作/事件叙事
  • 含文字视频 → 结构化内容报告

单一编码器的设计

常见多模态模型用分开的视觉塔 + 音频塔,Youtu-Parsing-Omni 改成一个统一的 Youtu-Omni-Encoder:

  • 模态-specific 的薄层 stem 把像素/log-mel 频谱图映射成 token
  • 所有 token 在同一个双向 Transformer 里处理,用 (t, h, w) 统一位置编码
  • 图片、音频帧、视频帧可以打包进同一次前向计算

视频处理时音视频融合层在每个音频到帧的时间边界打开新的注意力窗口,视频帧可以「看到」同一时刻的声音。

训练用了作者自己提出的 OSAD 流程:SFT → schema 路由 RLVR → OmniSchema-Aware On-Policy Distillation,不依赖单独的教师模型,靠模型自采样 + 准入门控迭代自蒸馏。


Benchmark(自报)

OmniDocBench v1.6(文档解析,自报开源 SOTA):

指标分数
整体96.96(对比 TeleOCR 96.91)
表格 TEDS96.79
公式 CDM96.80

OmniParsingBench(跨模态,最强开源权重):

类别分数
平均75.08(Gemini-3-Pro 77.44)
图表95.02(超 Gemini-3-Pro 92.79)
音频78.08(超 Gemini-3-Pro 76.74)
几何图74.33
自然图像62.84

以上均为作者自报,论文 arXiv 编号标注为 TBD,未经外部独立复现验证。


推理环境(门槛较高)

两套不兼容的 Python 环境都要配:

vLLM 服务(推荐,benchmark 结果从这套来):

pip install vllm==0.19.0 transformers==5.2.0
pip install vllm-plugin-vita-omni   # 自定义插件,不可缺
# ⚠️ 这套环境里不能装 peft
MODEL=tencent/Youtu-Parsing-Omni GPU=0 bash scripts/vllm.sh
# 启动 OpenAI 兼容服务 127.0.0.1:8000
python examples/infer_vllm.py --task document --media file.pdf

Transformers 推理:

pip install transformers==5.10.2
# trust_remote_code=True,建议用 --revision 锁定 commit

其他依赖:Python ≥ 3.10,CUDA GPU(BF16 权重),ffmpeg(音视频处理)。


License 注意

许可证结构类似 Apache-2.0,但附加了一条:“本软件不适用于欧盟境内使用”。如与标准 Apache-2.0 有冲突,以该条款为准。EU 用户需注意。


已知边界

  • 自报 benchmark,论文 arXiv 编号 TBD,尚无外部独立验证
  • 需要两套不兼容环境,设置复杂度较高
  • natural_image 任务在 OmniParsingBench 得分 62.84,低于其他模态
  • 88 stars,极早期,生产稳定性未经验证
  • License 含 EU 排除条款,不是标准开源

一句话说清楚

Youtu-Parsing-Omni 是腾讯 Youtu Lab 的 5B 全模态解析模型:一个统一编码器覆盖 8 种输入,所有任务输出统一 OmniSchema JSON,文档/图表/音频/视频的感知与认知结果结构一致。OmniDocBench 96.96(自报 SOTA),OmniParsingBench 75.08(最强开源权重)。推理依赖 vLLM 0.19.0 + 自定义插件,门槛不低。


88 stars,腾讯 Youtu Lab,License 含 EU 排除条款。开源仅供学习参考,EU 用户注意 License 条款,论文 arXiv 编号 TBD。


🇬🇧 English

Tencent Youtu-Parsing-Omni: One 5B Model to Parse Documents, Charts, Audio, and Video

Most parsing tools handle one modality: documents, images, or audio — specialized and siloed. Youtu-Parsing-Omni aims to do it in one: one 5B model, one unified JSON, eight input modalities.

GitHub: https://github.com/TencentCloudADP/youtu-parsing HuggingFace: https://huggingface.co/tencent/Youtu-Parsing-Omni

88 stars, Tencent Youtu Lab, published under TencentCloudADP.


Eight Modalities, One Output

The --task parameter selects the task:

TaskInput
documentDocument pages (PDF, scans)
natural_imagePhotographs or general images
graphics_chartBar/line/pie charts
graphics_flowchartFlowcharts and diagrams
graphics_geometricGeometry figures (math/science)
audioAudio clips
natural_videoVideo without heavy text overlay
textrich_videoText-rich video (slides, screencasts)

All tasks emit the same OmniSchema JSON — field population varies by task type. Fields: modality, subtype, elements (with type, content, bbox), reading_order, segments (timestamped for audio/video), global_description.

Bounding boxes use special tokens on a 0–1000 normalized grid: <box><x_85><y_275><x_1000><y_1000></box>


What the Perception Layer Outputs

InputPerception Output
DocumentLayout blocks + OCR + bbox + OTSL tables + LaTeX formulas + reading order
Chart/diagramMarkdown table (charts), Mermaid (flowcharts), point/arc/measurement (geometry)
AudioTimestamps + speaker ASR + non-vocal acoustic event labels
Natural videoTimestamps + scene descriptions + camera motion labels
Text-rich videoTimestamps + ASR + OCR + structured_report (Markdown)

Single Unified Encoder Design

Rather than separate vision and audio towers, Youtu-Parsing-Omni uses one unified Youtu-Omni-Encoder: modality-specific thin stems map pixels/log-mel spectrograms to tokens; all tokens process together in a shared bidirectional Transformer with (t, h, w) positional encoding. Images, audio chunks, and video frames can be packed into a single forward pass.

For video, fusion layers open a new attention window at each audio-to-frame temporal boundary, so video frames attend to the sound from the same moment.

Training pipeline: SFT → schema-routed RLVR → OmniSchema-Aware On-Policy Distillation (OSAD). No separate teacher model — the model self-samples with an admission gate (JSON validity, schema conformance, box syntax, repetition), then iteratively distills from itself.


Benchmarks (Self-Reported)

OmniDocBench v1.6 (document parsing, claimed best open-weight):

MetricScore
Overall96.96 (vs TeleOCR 96.91)
Table TEDS96.79
Formula CDM96.80

OmniParsingBench (cross-modal, best open-weight):

CategoryScore
Average75.08 (Gemini-3-Pro 77.44)
Chart95.02 (beats Gemini-3-Pro 92.79)
Audio78.08 (beats Gemini-3-Pro 76.74)
Geometry74.33
Natural image62.84

All numbers are self-reported. The paper arXiv ID is listed as TBD — no independent external validation found yet.


Inference Setup (Non-Trivial)

Two incompatible Python environments are required:

vLLM (recommended; benchmark results from this):

pip install vllm==0.19.0 transformers==5.2.0
pip install vllm-plugin-vita-omni  # custom plugin, required
# Do NOT install peft in this env
MODEL=tencent/Youtu-Parsing-Omni GPU=0 bash scripts/vllm.sh
python examples/infer_vllm.py --task document --media file.pdf

Transformers inference:

pip install transformers==5.10.2  # incompatible with vLLM env
# trust_remote_code=True, pin with --revision for safety

Requirements: Python ≥ 3.10, CUDA GPU (BF16 weights), ffmpeg for audio/video.


License Warning

The license is Apache-2.0 in structure, but adds one clause: “IS NOT INTENDED FOR USE WITHIN THE EUROPEAN UNION” — this clause prevails in case of conflict. EU users need to be aware.


Known Limits

  • Self-reported benchmarks; arXiv ID TBD; no independent external validation yet
  • Two incompatible environments make setup non-trivial
  • natural_image score (62.84) notably lower than other modalities
  • 88 stars, very early stage
  • License has EU exclusion — not standard open source

TL;DR

Youtu-Parsing-Omni is Tencent Youtu Lab’s 5B omni-modal parsing model: one unified encoder for 8 input types, all tasks emit a unified OmniSchema JSON. Self-reported SOTA on OmniDocBench (96.96) and best open-weight on OmniParsingBench (75.08 vs Gemini-3-Pro 77.44). Inference requires vLLM 0.19.0 + custom plugin; setup bar is not low. License excludes EU use.


88 stars, Tencent Youtu Lab, custom license (EU excluded). For reference only — EU users check the LICENSE carefully; arXiv ID TBD.

💬 评论与讨论

使用 GitHub 账号登录后发表评论

关于本站 · 免责声明

🍄 Mushroom Research Blog 是非营利、免费公开的个人科技观察博客与公众号 XStack18,不接受商业合作、不代表任何企业或机构立场,也不谋求商业利益。我们以个人视角客观中立地记录和分析 AI、Web3 等领域的最新模型发布与技术动态——不止转述新闻标题或二手信息,而是给出有独立思考的深入分析,希望帮更多人获得有价值的一手科技认知。

⚠️ 文中介绍的开源代码与模型,仅供学习交流与技术借鉴。它们大多仍处于早期阶段,有待进一步研究和验证,请勿直接用于工作或生产环境;如需采用,请先自行充分测试,并核实其许可证与安全性。
Open-source code and models featured here are shared for learning and reference only. Most are early-stage and still need further study and verification — please don't use them directly in your work or in production. Test them thoroughly and check their licenses and security first.

  1. 本站文章均为作者基于公开信息的个人研究与观点整理,不代表文中提及的任何公司、产品、模型的官方立场,未与其构成商业关联或合作关系。
  2. 科技行业信息更新极快,我们尽力保证内容准确、及时,但不对完整性、实时性做绝对保证,具体请以相关企业/项目官方公告为准。
  3. 文中引用的第三方商标、产品名称、图片、数据等版权归原权利人所有,我们会尽量注明来源;如你认为存在版权疑问或侵权,请通过下方邮箱联系我们,收到通知后会尽快核实处理(更正、加注来源或删除)。
  4. 文章内容仅为技术科普与个人观点,不构成投资、法律或其他专业建议,据此进行任何决策的后果需自行判断和承担。

📮 侵权 / 勘误 / 合作咨询:[email protected]