4 papers
Gaze Target Estimation Anywhere with Concepts
Xu Cao, Houze Yang, Vipin Gunda +5
Estimating human gaze targets from images in-the-wild is an important and formidable task. Existing approaches primarily employ brittle, multi-stage pipelines that require explicit…
MoXaRt: Audio-Visual Object-Guided Sound Interaction for XR
Tianyu Xu, Sieun Kim, Qianhui Zheng +6
In Extended Reality (XR), complex acoustic environments often overwhelm users, compromising both scene awareness and social engagement due to entangled sound sources. We introduce…
SPAN-Nav: Generalized Spatial Awareness for Versatile Vision-Language Navigation
Jiahang Liu, Tianyu Xu, Jiawei Chen +9
Recent embodied navigation approaches leveraging Vision-Language Models (VLMs) demonstrate strong generalization in versatile Vision-Language Navigation (VLN). However, reliable pa…
GCA-SUNet: A Gated Context-Aware Swin-UNet for Exemplar-Free Counting
Yuzhe Wu, Yipeng Xu, Tianyu Xu +3
Exemplar-Free Counting aims to count objects of interest without intensive annotations of objects or exemplars. To achieve this, we propose a Gated Context-Aware Swin-UNet (GCA-SUN…