2 papers
cs.CV2026
An LMM for Precisely Grounding Elements in Documents
Yijian Lu, Chuangxin Zhao, Kai Sun +3
Visual grounding in documents is a crucial ability for Large Multimodal Models (LMMs) in areas such as document understanding, deep research and document error detection. However,…
cs.CV2025
MMGeoLM: Hard Negative Contrastive Learning for Fine-Grained Geometric Understanding in Large Multimodal Models
Kai Sun, Yushi Bai, Zhen Yang +4
Large Multimodal Models (LMMs) typically build on ViTs (e.g., CLIP), yet their training with simple random in-batch negatives limits the ability to capture fine-grained visual diff…