7 papers
Zero-Shot Object Re-Identification in Egocentric Kitchen Videos via Multi-Stage SAM3 Feature Fusion
Dmytro Klepachevskyi, Alexander Wong, Sirisha Rambhatla +1
Object re-identification (ReID) in egocentric kitchen videos is challenging due to rapid viewpoint changes, frequent occlusions, cluttered scenes, and large intra-class appearance…
KitchenTwin: Semantically and Geometrically Grounded 3D Kitchen Digital Twins
Quanyun Wu, Kyle Gao, Daniel Long +3
Embodied AI training and evaluation require object-centric digital twin environments with accurate metric geometry and semantic grounding. Recent transformer-based feedforward reco…
Avatar4D: Synthesizing Domain-Specific 4D Humans for Real-World Pose Estimation
Jerrin Bright, Zhibo Wang, Dmytro Klepachevskyi +4
We present Avatar4D, a real-world transferable pipeline for generating customizable synthetic human motion datasets tailored to domain-specific applications. Unlike prior works, wh…
DreamPose3D: Hallucinative Diffusion with Prompt Learning for 3D Human Pose Estimation
Jerrin Bright, Yuhao Chen, John S. Zelek
Accurate 3D human pose estimation remains a critical yet unresolved challenge, requiring both temporal coherence across frames and fine-grained modeling of joint relationships. How…
Ice Hockey Puck Localization Using Contextual Cues
Liam Salass, Jerrin Bright, Amir Nazemi +3
Puck detection in ice hockey broadcast videos poses significant challenges due to the puck's small size, frequent occlusions, motion blur, broadcast artifacts, and scale inconsiste…
Gen4D: Synthesizing Humans and Scenes in the Wild
Jerrin Bright, Zhibo Wang, Yuhao Chen +3
Lack of input data for in-the-wild activities often results in low performance across various computer vision tasks. This challenge is particularly pronounced in uncommon human-cen…