5 papers
MotionVLA: Vision-Language-Action Model for Humanoid Motion
Nonghai Zhang, Siyu Zhai, Yanjun Li +5
Generating realistic humanoid motion from scene images and text involves both low-frequency pose semantics and high-frequency physical dynamics. However, many existing methods toke…
CMamba: Learned Image Compression with State Space Models
Zhuojie Wu, Heming Du, Shuyun Wang +4
Learned Image Compression (LIC) has explored various architectures, such as Convolutional Neural Networks (CNNs) and transformers, in modeling image content distributions in order…
Technical Report for ActivityNet Challenge 2022 -- Temporal Action Localization
Shimin Chen, Wei Li, Jianyang Gu +2
In the task of temporal action localization of ActivityNet-1.3 datasets, we propose to locate the temporal boundaries of each action and predict action class in untrimmed videos. W…
Technical Report for SoccerNet Challenge 2022 -- Replay Grounding Task
Shimin Chen, Wei Li, Jiaming Chu +3
In order to make full use of video information, we transform the replay grounding problem into a video action location problem. We apply a unified network Faster-TAD proposed by us…
The SkatingVerse Workshop & Challenge: Methods and Results
Jian Zhao, Lei Jin, Jianshu Li +16
The SkatingVerse Workshop & Challenge aims to encourage research in developing novel and accurate methods for human action understanding. The SkatingVerse dataset used for the Skat…