activity
20212024
most citedM2-CLIP: A Multimodal, Multi-task Adapting Framework for Video Action Recognition

4 citations · 17 across the 18 of their papers we have counts for

collaborators

18 papers

cs.HC20241 cited

MambaGesture: Enhancing Co-Speech Gesture Generation with Mamba and Disentangled Multi-Modality Fusion

Chencan Fu, Yabiao Wang, Jiangning Zhang +7

Co-speech gesture generation is crucial for producing synchronized and realistic human gestures that accompany speech, enhancing the animation of lifelike avatars in virtual enviro…

cs.CV20244 cited

M2-CLIP: A Multimodal, Multi-task Adapting Framework for Video Action Recognition

Mengmeng Wang, Jiazheng Xing, Boyuan Jiang +6

Recently, the rise of large-scale vision-language pretrained models like CLIP, coupled with the technology of Parameter-Efficient FineTuning (PEFT), has captured substantial attrac…

cs.CV20232 cited

A Generalist FaceX via Learning Unified Facial Representation

Yue Han, Jiangning Zhang, Junwei Zhu +7

This work presents FaceX framework, a novel facial generalist model capable of handling diverse facial tasks simultaneously. To achieve this goal, we initially formulate a unified…

cs.LG2023

Unified Data-Free Compression: Pruning and Quantization without Fine-Tuning

Shipeng Bai, Jun Chen, Xintian Shen +2

Structured pruning and quantization are promising approaches for reducing the inference time and memory footprint of neural networks. However, most existing methods require the ori…

cs.CV2023

High-fidelity Generalized Emotional Talking Face Generation with Multi-modal Emotion Space Learning

Chao Xu, Junwei Zhu, Jiangning Zhang +6

Recently, emotional talking face generation has received considerable attention. However, existing methods only adopt one-hot coding, image, or audio as emotion conditions, thus la…

cs.CV2023

Correlation Pyramid Network for 3D Single Object Tracking

Mengmeng Wang, Teli Ma, Xingxing Zuo +2

3D LiDAR-based single object tracking (SOT) has gained increasing attention as it plays a crucial role in 3D applications such as autonomous driving. The central problem is how to…