2 papers
cs.CV2025
TernaryCLIP: Efficiently Compressing Vision-Language Models with Ternary Weights and Distilled Knowledge
Shu-Hao Zhang, Wei-Cheng Tang, Chen Wu +5
Recent years have witnessed an increasing interest in image-text contrastive modeling, exemplified by models such as Contrastive Language-Image Pretraining (CLIP). In this paper, w…
cs.CV2025
Visual Position Prompt for MLLM based Visual Grounding
Wei Tang, Yanpeng Sun, Qinying Gu +1
Although Multimodal Large Language Models (MLLMs) excel at various image-related tasks, they encounter challenges in precisely aligning coordinates with spatial information within…