14 citations · 34 across the 8 of their papers we have counts for
8 papers
MMRA: A Benchmark for Evaluating Multi-Granularity and Multi-Image Relational Association Capabilities in Large Visual Language Models
Siwei Wu, Kang Zhu, Yu Bai +10
Given the remarkable success that large visual language models (LVLMs) have achieved in image perception tasks, the endeavor to make LVLMs perceive the world like humans is drawing…
ML-LMCL: Mutual Learning and Large-Margin Contrastive Learning for Improving ASR Robustness in Spoken Language Understanding
Xuxin Cheng, Bowen Cao, Qichen Ye +3
Spoken language understanding (SLU) is a fundamental task in the task-oriented dialogue systems. However, the inevitable errors from automatic speech recognition (ASR) usually impa…
Extreme Parkour with Legged Robots
Xuxin Cheng, Kexin Shi, Ananye Agarwal +1
Humans can perform parkour by traversing obstacles in a highly dynamic fashion requiring precise eye-muscle coordination and movement. Getting robots to do the same task requires o…
Legs as Manipulator: Pushing Quadrupedal Agility Beyond Locomotion
Xuxin Cheng, Ashish Kumar, Deepak Pathak
Locomotion has seen dramatic progress for walking or running across challenging terrains. However, robotic quadrupeds are still far behind their biological counterparts, such as do…
PoseRAC: Pose Saliency Transformer for Repetitive Action Counting
Ziyu Yao, Xuxin Cheng, Yuexian Zou
This paper presents a significant contribution to the field of repetitive action counting through the introduction of a new approach called Pose Saliency Representation. The propos…
SSVMR: Saliency-based Self-training for Video-Music Retrieval
Xuxin Cheng, Zhihong Zhu, Hongxiang Li +2
With the rise of short videos, the demand for selecting appropriate background music (BGM) for a video has increased significantly, video-music retrieval (VMR) task gradually draws…