120 citations · 157 across the 6 of their papers we have counts for
13 papers
CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation
Vishnu Sashank Dorbala, Gunnar Sigurdsson, Robinson Piramuthu +2
Household environments are visually diverse. Embodied agents performing Vision-and-Language Navigation (VLN) in the wild must be able to handle this diversity, while also following…
Video in 10 Bits: Few-Bit VideoQA for Efficiency and Privacy
Shiyuan Huang, Robinson Piramuthu, Shih-Fu Chang +1
In Video Question Answering (VideoQA), answering general questions about a video requires its visual information. Yet, video often contains redundant information irrelevant to the…
Self-Attentive 3D Human Pose and Shape Estimation from Videos
Yun-Chun Chen, Marco Piccirilli, Robinson Piramuthu +1
We consider the task of estimating 3D human pose and shape from videos. While existing frame-based approaches have made significant progress, these methods are independently applie…
Mixup-CAM: Weakly-supervised Semantic Segmentation via Uncertainty Regularization
Yu-Ting Chang, Qiaosong Wang, Wei-Chih Hung +3
Obtaining object response maps is one important step to achieve weakly-supervised semantic segmentation using image-level labels. However, existing methods rely on the classificati…
Weakly-Supervised Semantic Segmentation via Sub-category Exploration
Yu-Ting Chang, Qiaosong Wang, Wei-Chih Hung +3
Existing weakly-supervised semantic segmentation methods using image-level annotations typically rely on initial responses to locate object regions. However, such response maps gen…
Mobile Head Tracking for eCommerce and Beyond
Muratcan Cicek, Jinrong Xie, Qiaosong Wang +1
Shopping is difficult for people with motor impairments. This includes online shopping. Proprietary software can emulate mouse and keyboard via head tracking. However, such a solut…