2 papers
cs.CV2026
HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better
Gengluo Li, Xingyu Wan, Shangpin Peng +20
We present HunyuanOCR-1.5, a lightweight end-to-end OCR-specialized vision-language model. HunyuanOCR unifies document parsing, text spotting, information extraction, text-image tr…
cs.CV2026
RefineSVG: Visual Feedback-Driven Reinforcement Learning for Image-to-SVG Generation
Shaobo Liu, Feiqiao Mao, Shuaishuai Zhou +4
We propose RefineSVG, a single-step closed-loop visual feedback framework that enables multimodal large language models (MLLMs) to perform high-fidelity image-to-SVG generation thr…