3 papers
cs.LG2026
OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning
Mingxin Huang, Yongxin Shi, Dezhi Peng +3
Recent advancements in multimodal slow-thinking systems have demonstrated remarkable performance across various visual reasoning tasks. However, their capabilities in text-rich ima…
cs.CV2025
Privacy-Preserving Biometric Verification with Handwritten Random Digit String
Peirong Zhang, Yuliang Liu, Songxuan Lai +2
Handwriting verification has stood as a steadfast identity authentication method for decades. However, this technique risks potential privacy breaches due to the inclusion of perso…
cs.CV2024
DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming
Jiaxin Zhang, Wentao Yang, Songxuan Lai +2
Current multimodal large language models (MLLMs) face significant challenges in visual document understanding (VDU) tasks due to the high resolution, dense text, and complex layout…