3 papers
cs.CV2026
NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
Peng Cai, Zhaofan Zou, Shifa Liu +7
Document parsing aims to transform unstructured documents into structured and machine-readable representations. Recent advances in Vision-Language Models (VLMs) have significantly…
cs.CV2025
OmniEraser: Remove Objects and Their Effects in Images with Paired Video-Frame Data
Runpu Wei, Zijin Yin, Shuo Zhang +8
Inpainting algorithms have achieved remarkable progress in removing objects from images, yet still face two challenges: 1) struggle to handle the object's visual effects such as sh…
cs.CV2025
Detailed Object Description with Controllable Dimensions
Xinran Wang, Haiwen Zhang, Baoteng Li +5
Object description plays an important role for visually impaired individuals to understand and compare the differences between objects. Recent multimodal large language models(MLLM…