2 papers
cs.CV2026
Affogato: Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale
Junha Lee, Eunha Park, Chunghyun Park +2
Affordance grounding aims to localize where to interact with an object, a fundamental capability for embodied agents. Yet progress is bottlenecked by data: manual annotation is pro…
cs.CV2024
DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignment
Cijo Jose, Théo Moutakanni, Dahyun Kang +11
Self-supervised visual foundation models produce powerful embeddings that achieve remarkable performance on a wide range of downstream tasks. However, unlike vision-language models…