As of Oct 9, Tencent Youtu Lab released Youtu-Parsing-Omni, a 5B omni-modal parsing model with a single shared encoder that turns document pages, natural images, charts, flowcharts, geometry figures, audio and video into one structured JSON schema.
Tencent Youtu Lab released Youtu-Parsing-Omni, a 5B omni-modal parsing model with a single shared encoder that turns document pages, natural images, charts, flowcharts, geometry figures, audio and video into one structured JSON schema. Tencent reports 96.96 overall on OmniDocBench v1.6 and 75.08 on OmniParsingBench, second only to Gemini-3-Pro, and ships weights, a technical report and a vLLM plugin.
Items older than 72 hours never go in Top stories and show their original date.