8 papers
Embracing Aleatoric Uncertainty: Generating Diverse 3D Human Motion
Zheng Qin, Yabing Wang, Minghui Yang +3
Generating 3D human motions from text is a challenging yet valuable task. The key aspects of this task are ensuring text-motion consistency and achieving generation diversity. Alth…
HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses through Reasoning MLLMs
Zheng Qin, Ruobing Zheng, Yabing Wang +4
While Multimodal Large Language Models (MLLMs) show immense promise for achieving truly human-like interactions, progress is hindered by the lack of fine-grained evaluation framewo…
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning
Yizhe Li, Sanping Zhou, Zheng Qin +1
Dense video captioning is a challenging task that aims to localize and caption multiple events in an untrimmed video. Recent studies mainly follow the transformer-based architectur…
From Mapping to Composing: A Two-Stage Framework for Zero-shot Composed Image Retrieval
Yabing Wang, Zhuotao Tian, Qingpei Guo +4
Composed Image Retrieval (CIR) is a challenging multimodal task that retrieves a target image based on a reference image and accompanying modification text. Due to the high cost of…
RSRNav: Reasoning Spatial Relationship for Image-Goal Navigation
Zheng Qin, Le Wang, Yabing Wang +3
Recent image-goal navigation (ImageNav) methods learn a perception-action policy by separately capturing semantic features of the goal and egocentric images, then passing them to a…
Versatile Multimodal Controls for Expressive Talking Human Animation
Zheng Qin, Ruobing Zheng, Yabing Wang +5
In filmmaking, directors typically allow actors to perform freely based on the script before providing specific guidance on how to present key actions. AI-generated content faces s…