6 papers · 1 filter
AdaCount: Training-Free Similarity-Guided Spatial and Feature Adaptation for Zero-Shot Object Counting
Muhammad Ibraheem Siddiqui, Muhammad Haris Khan
Zero-shot object counting (ZOC) aims to count instances of arbitrary object categories specified only through textual prompts. Recent training-free approaches leverage foundation m…
Not All Modalities Are Equal: Instruction-Aware Gating for Multimodal Videos
Bonan Ding, Umair Nawaz, Ufaq Khan +5
Pre-trained video large language models excel at visual reasoning. However, they struggle when videos arrive with auxiliary streams, such as audio, depth map, or dense temporal evi…
FCC: Fully Connected Correlation for One-Shot Segmentation
Seonghyeon Moon, Haein Kong, Muhammad Haris Khan +2
Few-shot segmentation (FSS) aims to segment the target object in a query image using only a small set of support images and masks. Therefore, having strong prior information for th…
Judging from Support-set: A New Way to Utilize Few-Shot Segmentation for Segmentation Refinement Process
Seonghyeon Moon, Qingze, Liu +2
Segmentation refinement aims to enhance the initial coarse masks generated by segmentation algorithms. The refined masks are expected to capture more details and better contours of…
PerSense: Training-Free Personalized Instance Segmentation in Dense Images
Muhammad Ibraheem Siddiqui, Muhammad Umer Sheikh, Hassan Abid +1
The emergence of foundational models has significantly advanced segmentation approaches. However, challenges still remain in dense scenarios, where occlusions, scale variations, an…
Why Not Use Your Textbook? Knowledge-Enhanced Procedure Planning of Instructional Videos
Kumaranage Ravindu Yasas Nagasinghe, Honglu Zhou, Malitha Gunawardhana +3
In this paper, we explore the capability of an agent to construct a logical sequence of action steps, thereby assembling a strategic procedural plan. This plan is crucial for navig…