5 papers
SeGDeP: Semantic- and Geometric-Aware Decoupled Prompts for Reasoning Segmentation
Linnan Zhao, Xu Liu, Lingling Li +3
Reasoning segmentation converts an implicit linguistic conclusion into a precise mask, requiring both semantic identification and spatial grounding. Existing MLLM-segmenter interfa…
OCR-EDR: Rendering-Aware Diagnosis and Repair for Closed-Loop OCR Improvement
Linnan Zhao, Kang Liu, Hao Yu +3
Although document OCR systems perform increasingly well on routine documents, complex formulas, structured text, and long-tail formats remain error-prone. OCR predictions may omit…
PaDoc: Layout-Grounded Parallel Decoding for Document Parsing
Hao Yu, Jiabo Zhan, Kang Liu +8
End-to-end document parsers provide a unified interface, but serialize page layouts and regional contents into one autoregressive sequence. This formulation forces independent regi…
PVUW 2025 Challenge Report: Advances in Pixel-level Understanding of Complex Videos in the Wild
Henghui Ding, Chang Liu, Nikhila Ravi +33
This report provides a comprehensive overview of the 4th Pixel-level Video Understanding in the Wild (PVUW) Challenge, held in conjunction with CVPR 2025. It summarizes the challen…
MASSeg : 2nd Technical Report for 4th PVUW MOSE Track
Xuqiang Cao, Linnan Zhao, Jiaxuan Zhao +3
Complex video object segmentation continues to face significant challenges in small object recognition, occlusion handling, and dynamic scene modeling. This report presents our sol…