3 papers
cs.CV2025
One Last Attention for Your Vision-Language Model
Liang Chen, Ghazi Shazan Ahmad, Tianjun Yao +2
Pretrained vision-language models (VLMs), such as CLIP, achieve remarkable zero-shot performance, yet their downstream potential hinges on effective fine-tuning. Most adaptation me…
cs.CV2025
VideoMolmo: Spatio-Temporal Grounding Meets Pointing
Ghazi Shazan Ahmad, Ahmed Heakl, Hanan Gani +4
Spatio-temporal localization is vital for precise interactions across diverse domains, from biological research to autonomous navigation and interactive interfaces. Current video-b…
cs.AI2024
ScaleViz: Scaling Visualization Recommendation Models on Large Data
Ghazi Shazan Ahmad, Shubham Agarwal, Subrata Mitra +4
Automated visualization recommendations (vis-rec) help users to derive crucial insights from new datasets. Typically, such automated vis-rec models first calculate a large number o…