7 papers
HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses through Reasoning MLLMs
Zheng Qin, Ruobing Zheng, Yabing Wang +4
While Multimodal Large Language Models (MLLMs) show immense promise for achieving truly human-like interactions, progress is hindered by the lack of fine-grained evaluation framewo…
Embracing Aleatoric Uncertainty: Generating Diverse 3D Human Motion
Zheng Qin, Yabing Wang, Minghui Yang +3
Generating 3D human motions from text is a challenging yet valuable task. The key aspects of this task are ensuring text-motion consistency and achieving generation diversity. Alth…
RSRNav: Reasoning Spatial Relationship for Image-Goal Navigation
Zheng Qin, Le Wang, Yabing Wang +3
Recent image-goal navigation (ImageNav) methods learn a perception-action policy by separately capturing semantic features of the goal and egocentric images, then passing them to a…
From Mapping to Composing: A Two-Stage Framework for Zero-shot Composed Image Retrieval
Yabing Wang, Zhuotao Tian, Qingpei Guo +4
Composed Image Retrieval (CIR) is a challenging multimodal task that retrieves a target image based on a reference image and accompanying modification text. Due to the high cost of…
Moment Quantization for Video Temporal Grounding
Xiaolong Sun, Le Wang, Sanping Zhou +5
Video temporal grounding is a critical video understanding task, which aims to localize moments relevant to a language description. The challenge of this task lies in distinguishin…
Referencing Where to Focus: Improving VisualGrounding with Referential Query
Yabing Wang, Zhuotao Tian, Qingpei Guo +4
Visual Grounding aims to localize the referring object in an image given a natural language expression. Recent advancements in DETR-based visual grounding methods have attracted co…