4 citations · 6 across the 6 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
OCR-Agent: Agentic OCR with Capability and Memory Reflection
Shimin Wen, Zeyu Zhang, Xingdou Bian +5
Large Vision-Language Models (VLMs) have demonstrated significant potential on complex visual understanding tasks through iterative optimization methods.However, these models gener…
cs.CV2026
OmniOCR: Generalist OCR for Ethnic Minority Languages
Bonan Liu, Zeyu Zhang, Bingbing Meng +5
Optical character recognition (OCR) has advanced rapidly with deep learning and multimodal models, yet most methods focus on well-resourced scripts such as Latin and Chinese. Ethni…
cs.CV2025
DOEI: Dual Optimization of Embedding Information for Attention-Enhanced Class Activation Maps
Hongjie Zhu, Zeyu Zhang, Guansong Pang +6
Weakly supervised semantic segmentation (WSSS) typically utilizes limited semantic annotations to obtain initial Class Activation Maps (CAMs). However, due to the inadequate coupli…