collaborators

7 papers

cs.CV2025

HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses through Reasoning MLLMs

Zheng Qin, Ruobing Zheng, Yabing Wang +4

While Multimodal Large Language Models (MLLMs) show immense promise for achieving truly human-like interactions, progress is hindered by the lack of fine-grained evaluation framewo…

cs.CV2025

Embracing Aleatoric Uncertainty: Generating Diverse 3D Human Motion

Zheng Qin, Yabing Wang, Minghui Yang +3

Generating 3D human motions from text is a challenging yet valuable task. The key aspects of this task are ensuring text-motion consistency and achieving generation diversity. Alth…

cs.CV2025

RSRNav: Reasoning Spatial Relationship for Image-Goal Navigation

Zheng Qin, Le Wang, Yabing Wang +3

Recent image-goal navigation (ImageNav) methods learn a perception-action policy by separately capturing semantic features of the goal and egocentric images, then passing them to a…

cs.CV2025

From Mapping to Composing: A Two-Stage Framework for Zero-shot Composed Image Retrieval

Yabing Wang, Zhuotao Tian, Qingpei Guo +4

Composed Image Retrieval (CIR) is a challenging multimodal task that retrieves a target image based on a reference image and accompanying modification text. Due to the high cost of…

cs.CV2025

Moment Quantization for Video Temporal Grounding

Xiaolong Sun, Le Wang, Sanping Zhou +5

Video temporal grounding is a critical video understanding task, which aims to localize moments relevant to a language description. The challenge of this task lies in distinguishin…

cs.CV2024

Referencing Where to Focus: Improving VisualGrounding with Referential Query

Yabing Wang, Zhuotao Tian, Qingpei Guo +4

Visual Grounding aims to localize the referring object in an image given a natural language expression. Recent advancements in DETR-based visual grounding methods have attracted co…