4 papers · 1 filter
VESSA: Video-based objEct-centric Self-Supervised Adaptation for Visual Foundation Models
Jesimon Barreto, Carlos Caetano, André Araujo +1
Foundation models have advanced computer vision by enabling strong performance across diverse tasks through large-scale pretraining and supervised fine-tuning. However, they may un…
Infusing fine-grained visual knowledge to Vision-Language Models
Nikolaos-Antonios Ypsilantis, Kaifeng Chen, André Araujo +1
Large-scale contrastive pre-training produces powerful Vision-and-Language Models (VLMs) capable of generating representations (embeddings) effective for a wide variety of visual a…
UDON: Universal Dynamic Online distillatioN for generic image representations
Nikolaos-Antonios Ypsilantis, Kaifeng Chen, André Araujo +1
Universal image representations are critical in enabling real-world fine-grained and instance-level recognition applications, where objects and entities from any domain must be ide…
HAMMR: HierArchical MultiModal React agents for generic VQA
Lluis Castrejon, Thomas Mensink, Howard Zhou +3
Combining Large Language Models (LLMs) with external specialized tools (LLMs+tools) is a recent paradigm to solve multimodal tasks such as Visual Question Answering (VQA). While th…