Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
VESSA: Video-based objEct-centric Self-Supervised Adaptation for Visual Foundation Models
Jesimon Barreto, Carlos Caetano, André Araujo +1
Foundation models have advanced computer vision by enabling strong performance across diverse tasks through large-scale pretraining and supervised fine-tuning. However, they may un…
cs.CV2025
Infusing fine-grained visual knowledge to Vision-Language Models
Nikolaos-Antonios Ypsilantis, Kaifeng Chen, André Araujo +1
Large-scale contrastive pre-training produces powerful Vision-and-Language Models (VLMs) capable of generating representations (embeddings) effective for a wide variety of visual a…