collaborators

5 papers

cs.CV2026

VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA

Haibin He, Maoyuan Ye, Jing Zhang +2

Video text-based visual question answering (Video TextVQA) aims to answer questions by reasoning over visual textual content appearing in videos. Despite the strong multimodal vide…

cs.CV2025

GoMatching++: Parameter- and Data-Efficient Arbitrary-Shaped Video Text Spotting and Benchmarking

Haibin He, Jing Zhang, Maoyuan Ye +3

Video text spotting (VTS) extends image text spotting (ITS) by adding text tracking, significantly increasing task complexity. Despite progress in VTS, existing methods still fall…

cs.CV2025

Adapting Segment Anything Model for Power Transmission Corridor Hazard Segmentation

Hang Chen, Maoyuan Ye, Peng Yang +3

Power transmission corridor hazard segmentation (PTCHS) aims to separate transmission equipment and surrounding hazards from complex background, conveying great significance to mai…

cs.CV2025

SFA: Scan, Focus, and Amplify toward Guidance-aware Answering for Video TextVQA

Haibin He, Qihuang Zhong, Juhua Liu +3

Video text-based visual question answering (Video TextVQA) task aims to answer questions about videos by leveraging the visual text appearing within the videos. This task poses sig…

cs.CV2025

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues?

Haibin He, Maoyuan Ye, Jing Zhang +4

Large Multimodal Models (LMMs) have become increasingly versatile, accompanied by impressive Optical Character Recognition (OCR) related capabilities. Existing OCR-related benchmar…