3 citations · 3 across the 7 of their papers we have counts for
5 papers · 1 filter
OCR-Agent: Agentic OCR with Capability and Memory Reflection
Shimin Wen, Zeyu Zhang, Xingdou Bian +5
Large Vision-Language Models (VLMs) have demonstrated significant potential on complex visual understanding tasks through iterative optimization methods.However, these models gener…
OmniOCR: Generalist OCR for Ethnic Minority Languages
Bonan Liu, Zeyu Zhang, Bingbing Meng +5
Optical character recognition (OCR) has advanced rapidly with deep learning and multimodal models, yet most methods focus on well-resourced scripts such as Latin and Chinese. Ethni…
SSS: Semi-Supervised SAM-2 with Efficient Prompting for Medical Imaging Segmentation
Hongjie Zhu, Xiwei Liu, Rundong Xue +5
In the era of information explosion, efficiently leveraging large-scale unlabeled data while minimizing the reliance on high-quality pixel-level annotations remains a critical chal…
DOEI: Dual Optimization of Embedding Information for Attention-Enhanced Class Activation Maps
Hongjie Zhu, Zeyu Zhang, Guansong Pang +6
Weakly supervised semantic segmentation (WSSS) typically utilizes limited semantic annotations to obtain initial Class Activation Maps (CAMs). However, due to the inadequate coupli…
SegStitch: Multidimensional Transformer for Robust and Efficient Medical Imaging Segmentation
Shengbo Tan, Zeyu Zhang, Ying Cai +5
Medical imaging segmentation plays a significant role in the automatic recognition and analysis of lesions. State-of-the-art methods, particularly those utilizing transformers, hav…