← r/LocalLLaMA
▲
24
 
1👁
r/LocalLLaMA · u/jacek2023 · 4h ago

tencent/Youtu-Parsing-Omni · Hugging Face

Youtu-Parsing-Omni is a compact (5B) omni-modal parsing model. Given a single input — a document page, a natural image, a chart / flowchart, a geometry figure, an audio clip or an audio-visual video — it produces one structured JSON envelope that covers both perception (layout elements, text, tables, formulas, bounding boxes, timestamps, ASR, OCR, acoustic events, camera motion) and cognition (captions, narratives, reports). The output family is selected by the task prompt (--task in the examples, keys of prompts/youtu_parsing_omni.json).

|Input|--task|modality / subtype|Key contents|
|:-|:-|:-|:-|
|Document page|document|image / document|layout elements with bbox, text / LaTeX / OTSL tables / Markdown charts / Mermaid flowcharts, reading order|
|Natural image|natural_image|image / natural_image|entities and text with bbox, tags, captions, global description|
|Chart|graphics_chart|image / document|one chart element: Markdown table, notes, caption|
|Flowchart|graphics_flowchart|image / document|one flowchart element: Mermaid, caption|
|Geometry figure|graphics_geometric|image / document|one geometric element: points, lines, arcs, shapes, geometric relations and measurements|
|Audio|audio|audio / –|vocal / non-vocal segments with timestamps, speakers, ASR, timbre / scene captions, acoustic events|
|Natural video|natural_video|video / natural_video|temporal segments with visual elements, actions, interactions, camera motion, audio track|
|Text-rich video|textrich_video|video / text_rich_video|segments with OCR + ASR and a Markdown structured_report of the whole video|

Highlights (see the technical report for details):

  • Unified schema – one JSON envelope for seven parsing families, driven by the task prompt.
  • Omni encoder – image, audio and interleaved audio-visual video inputs (frames + audio track) in a single model.
  • Strong results at a small size – state-of-the-art on OmniDocBench v1.6 (96.96 Overall), best open-weight model on OmniParsingBench (75.08 Avg., second only to Gemini-3-Pro), and competitive with specialized models on chemical-structure (ChemOCR) and music-score (PDMX-Synth) recognition (results).
  • Easy to serve – a vLLM plugin, pinned serving settings, task prompts and inference examples are included.
posted Fri, 09 Oct 2026 12:03:32 GMTseen 1 time
open on reddit ↗ 💬 5