activity
20222024
most citedM3ST: Mix at Three Levels for Speech Translation

14 citations · 34 across the 8 of their papers we have counts for

collaborators

8 papers

cs.CV2024

MMRA: A Benchmark for Evaluating Multi-Granularity and Multi-Image Relational Association Capabilities in Large Visual Language Models

Siwei Wu, Kang Zhu, Yu Bai +10

Given the remarkable success that large visual language models (LVLMs) have achieved in image perception tasks, the endeavor to make LVLMs perceive the world like humans is drawing…

cs.CL2023

ML-LMCL: Mutual Learning and Large-Margin Contrastive Learning for Improving ASR Robustness in Spoken Language Understanding

Xuxin Cheng, Bowen Cao, Qichen Ye +3

Spoken language understanding (SLU) is a fundamental task in the task-oriented dialogue systems. However, the inevitable errors from automatic speech recognition (ASR) usually impa…

cs.RO20233 cited

Extreme Parkour with Legged Robots

Xuxin Cheng, Kexin Shi, Ananye Agarwal +1

Humans can perform parkour by traversing obstacles in a highly dynamic fashion requiring precise eye-muscle coordination and movement. Getting robots to do the same task requires o…

cs.RO20232 cited

Legs as Manipulator: Pushing Quadrupedal Agility Beyond Locomotion

Xuxin Cheng, Ashish Kumar, Deepak Pathak

Locomotion has seen dramatic progress for walking or running across challenging terrains. However, robotic quadrupeds are still far behind their biological counterparts, such as do…

cs.CV20238 cited

PoseRAC: Pose Saliency Transformer for Repetitive Action Counting

Ziyu Yao, Xuxin Cheng, Yuexian Zou

This paper presents a significant contribution to the field of repetitive action counting through the introduction of a new approach called Pose Saliency Representation. The propos…

cs.MM20231 cited

SSVMR: Saliency-based Self-training for Video-Music Retrieval

Xuxin Cheng, Zhihong Zhu, Hongxiang Li +2

With the rise of short videos, the demand for selecting appropriate background music (BGM) for a video has increased significantly, video-music retrieval (VMR) task gradually draws…