7 citations · 8 across the 16 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
ABE-CLIP: Training-Free Attribute Binding Enhancement for Compositional Image-Text Matching
Qi Zhang, Yuxu Chen, Lei Deng +1
Contrastive Language-Image Pretraining (CLIP) has achieved remarkable performance in various multimodal tasks. However, it still struggles with compositional image-text matching, p…
cs.CV2025
MonkeyOCR v1.5 Technical Report: Unlocking Robust Document Parsing for Complex Patterns
Jiarui Zhang, Yuliang Liu, Zijun Wu +17
Document parsing is a core task in document intelligence, supporting applications such as information extraction, retrieval-augmented generation, and automated document analysis. H…