6 papers · 1 filter
PinPoint: Prompting with Informative Interior Points
Pouya Sadeghi, Shawn He, Pedro Pablo Guerrero Vela +3
Modern referring image segmentation pipelines couple a vision-language model (VLM) for grounding with a promptable segmenter such as the Segment Anything Model (SAM) for mask gener…
Zero-Shot Object Re-Identification in Egocentric Kitchen Videos via Multi-Stage SAM3 Feature Fusion
Dmytro Klepachevskyi, Alexander Wong, Sirisha Rambhatla +1
Object re-identification (ReID) in egocentric kitchen videos is challenging due to rapid viewpoint changes, frequent occlusions, cluttered scenes, and large intra-class appearance…
Avatar4D: Synthesizing Domain-Specific 4D Humans for Real-World Pose Estimation
Jerrin Bright, Zhibo Wang, Dmytro Klepachevskyi +4
We present Avatar4D, a real-world transferable pipeline for generating customizable synthetic human motion datasets tailored to domain-specific applications. Unlike prior works, wh…
LOCATEdit: Graph Laplacian Optimized Cross Attention for Localized Text-Guided Image Editing
Achint Soni, Meet Soni, Sirisha Rambhatla
Text-guided image editing aims to modify specific regions of an image according to natural language instructions while maintaining the general structure and the background fidelity…
LangDA: Building Context-Awareness via Language for Domain Adaptive Semantic Segmentation
Chang Liu, Bavesh Balaji, Saad Hossain +5
Unsupervised domain adaptation for semantic segmentation (DASS) aims to transfer knowledge from a label-rich source domain to a target domain with no labels. Two key approaches in…
Domain-Guided Masked Autoencoders for Unique Player Identification
Bavesh Balaji, Jerrin Bright, Sirisha Rambhatla +4
Unique player identification is a fundamental module in vision-driven sports analytics. Identifying players from broadcast videos can aid with various downstream tasks such as play…