4 papers
Affogato: Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale
Junha Lee, Eunha Park, Chunghyun Park +2
Affordance grounding aims to localize where to interact with an object, a fundamental capability for embodied agents. Yet progress is bottlenecked by data: manual annotation is pro…
Exploring High-Order Self-Similarity for Video Understanding
Manjin Kim, Heeseung Kwon, Karteek Alahari +1
Space-time self-similarity (STSS), which captures visual correspondences across frames, provides an effective way to represent temporal dynamics for video understanding. In this wo…
Few-Shot Pattern Detection via Template Matching and Regression
Eunchan Jo, Dahyun Kang, Sanghyun Kim +2
We address the problem of few-shot pattern detection, which aims to detect all instances of a given pattern, typically represented by a few exemplars, from an input image. Although…
Memory-Modular Classification: Learning to Generalize with Memory Replacement
Dahyun Kang, Ahmet Iscen, Eunchan Jo +3
We propose a novel memory-modular learner for image classification that separates knowledge memorization from reasoning. Our model enables effective generalization to new classes b…