document parsing 1end-to-end OCR 1markdown generation 1reinforcement learning 1synthetic data augmentation 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.CV2026
OvisOCR2 Technical Report
Shiyin Lu, Yinglun Li, Yu Xia +10
OvisOCR2 is a 0.8 B parameter end‑to‑end model that converts document page images into Markdown, handling text, formulas, tables, and visual regions, and achieves state‑of‑the‑art…
cs.CV2025
Ovis2.5 Technical Report
Shiyin Lu, Yang Li, Yu Xia +39
We present Ovis2.5, a successor to Ovis2 designed for native-resolution visual perception and strong multimodal reasoning. Ovis2.5 integrates a native-resolution vision transformer…
cs.CV2024
UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation
Lunhao Duan, Shanshan Zhao, Wenjun Yan +7
Recently, text-to-image generation models have achieved remarkable advancements, particularly with diffusion models facilitating high-quality image synthesis from textual descripti…