47 citations · 67 across the 29 of their papers we have counts for
10 papers · 1 filter
MimicParts: Part-aware Style Injection for Speech-Driven 3D Motion Generation
Lianlian Liu, YongKang He, Zhaojie Chu +2
Generating stylized 3D human motion from speech signals presents substantial challenges, primarily due to the intricate and fine-grained relationships among speech signals, individ…
VideoCoT: A Video Chain-of-Thought Dataset with Active Annotation Tool
Yan Wang, Yawen Zeng, Jingsheng Zheng +3
Multimodal large language models (MLLMs) are flourishing, but mainly focus on images with less attention than videos, especially in sub-fields such as prompt engineering, video cha…
Towards Zero-shot Human-Object Interaction Detection via Vision-Language Integration
Weiying Xue, Qi Liu, Qiwei Xiong +4
Human-object interaction (HOI) detection aims to locate human-object pairs and identify their interaction categories in images. Most existing methods primarily focus on supervised…
PointCore: Efficient Unsupervised Point Cloud Anomaly Detector Using Local-Global Features
Baozhu Zhao, Qiwei Xiong, Xiaohan Zhang +4
Three-dimensional point cloud anomaly detection that aims to detect anomaly data points from a training set serves as the foundation for a variety of applications, including indust…
CorrTalk: Correlation Between Hierarchical Speech and Facial Activity Variances for 3D Animation
Zhaojie Chu, Kailing Guo, Xiaofen Xing +3
Speech-driven 3D facial animation is a challenging cross-modal task that has attracted growing research interest. During speaking activities, the mouth displays strong motions, whi…
LAPP: Layer Adaptive Progressive Pruning for Compressing CNNs from Scratch
Pucheng Zhai, Kailing Guo, Fang Liu +2
Structured pruning is a commonly used convolutional neural network (CNN) compression approach. Pruning rate setting is a fundamental problem in structured pruning. Most existing wo…