5 papers
From Instruction to Event: Sound-Triggered Mobile Manipulation
Hao Ju, Shaofei Huang, Hongyu Li +4
Current mobile manipulation research predominantly follows an instruction-driven paradigm, where agents rely on predefined textual commands to execute tasks. However, this setting…
How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
Songsong Yu, Yuxin Chen, Hao Ju +15
Visual Spatial Reasoning (VSR) is a core human cognitive ability and a critical requirement for advancing embodied intelligence and autonomous systems. Despite recent progress in V…
When Words Smile: Generating Diverse Emotional Facial Expressions from Text
Haidong Xu, Meishan Zhang, Hao Ju +4
Enabling digital humans to express rich emotions has significant applications in dialogue systems, gaming, and other interactive scenarios. While recent advances in talking head sy…
AnomalyLMM: Bridging Generative Knowledge and Discriminative Retrieval for Text-Based Person Anomaly Search
Hao Ju, Hu Zhang, Zhedong Zheng
With growing public safety demands, text-based person anomaly search has emerged as a critical task, aiming to retrieve individuals with abnormal behaviors via natural language des…
Video2BEV: Transforming Drone Videos to BEVs for Video-based Geo-localization
Hao Ju, Shaofei Huang, Si Liu +1
Existing approaches to drone visual geo-localization predominantly adopt the image-based setting, where a single drone-view snapshot is matched with images from other platforms. Su…