2 papers
cs.CV2026
RefineSVG: Visual Feedback-Driven Reinforcement Learning for Image-to-SVG Generation
Shaobo Liu, Feiqiao Mao, Shuaishuai Zhou +4
RefineSVG introduces a closed-loop visual feedback system that lets large multimodal language models iteratively correct SVG code by rendering the output, comparing it to the targe…
cs.CV2026
HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better
Gengluo Li, Xingyu Wan, Shangpin Peng +20
We present HunyuanOCR-1.5, a lightweight end-to-end OCR-specialized vision-language model. HunyuanOCR unifies document parsing, text spotting, information extraction, text-image tr…