Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
OmniHandwritingOCR: A Diagnostic Benchmark for Evaluating Multimodal LLMs in Handwritten OCR Scenarios
Zinuo Guo, Min Zhang, Bo Jiang
Multimodal large language models (MLLMs) are increasingly used as OCR systems in document and knowledge-processing pipelines, but their ability to faithfully read real handwriting…
cs.CV2026
Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models
Xinpeng Dong, Min Zhang, Kairong Han +3
In recent years, multimodal large language models (MLLMs) have achieved remarkable progress, primarily attributed to effective paradigms for integrating visual and textual informat…