6 papers
AdaCount: Training-Free Similarity-Guided Spatial and Feature Adaptation for Zero-Shot Object Counting
Muhammad Ibraheem Siddiqui, Muhammad Haris Khan
Zero-shot object counting (ZOC) aims to count instances of arbitrary object categories specified only through textual prompts. Recent training-free approaches leverage foundation m…
Not All Modalities Are Equal: Instruction-Aware Gating for Multimodal Videos
Bonan Ding, Umair Nawaz, Ufaq Khan +5
Pre-trained video large language models excel at visual reasoning. However, they struggle when videos arrive with auxiliary streams, such as audio, depth map, or dense temporal evi…
MASt3R-Nav: WayPixel Navigation in Relative 3D Maps
Vansh Garg, Rohit Jayanti, Krish Pandya +5
Visual navigation ability is strongly tied to its underlying representation of the world. Unlike classical 3D maps that require globally-consistent geometry, image- or object-relat…
FCC: Fully Connected Correlation for One-Shot Segmentation
Seonghyeon Moon, Haein Kong, Muhammad Haris Khan +2
Few-shot segmentation (FSS) aims to segment the target object in a query image using only a small set of support images and masks. Therefore, having strong prior information for th…
PerSense: Training-Free Personalized Instance Segmentation in Dense Images
Muhammad Ibraheem Siddiqui, Muhammad Umer Sheikh, Hassan Abid +1
The emergence of foundational models has significantly advanced segmentation approaches. However, challenges still remain in dense scenarios, where occlusions, scale variations, an…
Judging from Support-set: A New Way to Utilize Few-Shot Segmentation for Segmentation Refinement Process
Seonghyeon Moon, Qingze, Liu +2
Segmentation refinement aims to enhance the initial coarse masks generated by segmentation algorithms. The refined masks are expected to capture more details and better contours of…