1 citations · 1 across the 3 of their papers we have counts for
4 papers
ActiveSAM: Image-Conditional Class Pruning for Fast and Accurate Open-Vocabulary Segmentation
Tran Dinh Tien, Zhiqiang Shen
Segment Anything Model 3 (SAM 3) provides a strong frozen backbone for concept-prompted segmentation, but applying it directly to open-vocabulary semantic segmentation (OVSS) is in…
One Last Attention for Your Vision-Language Model
Liang Chen, Ghazi Shazan Ahmad, Tianjun Yao +2
Pretrained vision-language models (VLMs), such as CLIP, achieve remarkable zero-shot performance, yet their downstream potential hinges on effective fine-tuning. Most adaptation me…
VideoMolmo: Spatio-Temporal Grounding Meets Pointing
Ghazi Shazan Ahmad, Ahmed Heakl, Hanan Gani +4
Spatio-temporal localization is vital for precise interactions across diverse domains, from biological research to autonomous navigation and interactive interfaces. Current video-b…
LFME: A Simple Framework for Learning from Multiple Experts in Domain Generalization
Liang Chen, Yong Zhang, Yibing Song +2
Domain generalization (DG) methods aim to maintain good performance in an unseen target domain by using training data from multiple source domains. While success on certain occasio…