4 citations · 9 across the 13 of their papers we have counts for
13 papers
FacialPulse: An Efficient RNN-based Depression Detection via Temporal Facial Landmarks
Ruiqi Wang, Jinyang Huang, Jie Zhang +6
Depression is a prevalent mental health disorder that significantly impacts individuals' lives and well-being. Early detection and intervention are crucial for effective treatment…
Training-Free Large Model Priors for Multiple-in-One Image Restoration
Xuanhua He, Lang Li, Yingying Wang +7
Image restoration aims to reconstruct the latent clear images from their degraded versions. Despite the notable achievement, existing methods predominantly focus on handling specif…
LDM-SVC: Latent Diffusion Model Based Zero-Shot Any-to-Any Singing Voice Conversion with Singer Guidance
Shihao Chen, Yu Gu, Jie Zhang +4
Any-to-any singing voice conversion (SVC) is an interesting audio editing technique, aiming to convert the singing voice of one singer into that of another, given only a few second…
Multi-granularity Correspondence Learning from Long-term Noisy Videos
Yijie Lin, Jie Zhang, Zhenyu Huang +3
Existing video-language studies mainly focus on learning short video clips, leaving long-term temporal dependencies rarely explored due to over-high computational cost of modeling…
CBVS: A Large-Scale Chinese Image-Text Benchmark for Real-World Short Video Search Scenarios
Xiangshuo Qiao, Xianxin Li, Xiaozhe Qu +5
Vision-Language Models pre-trained on large-scale image-text datasets have shown superior performance in downstream tasks such as image retrieval. Most of the images for pre-traini…
Modelling and Performance Analysis of the Over-the-Air Computing in Cellular IoT Networks
Ying Dong, Haonan Hu, Qiaoshou Liu +3
Ultra-fast wireless data aggregation (WDA) of distributed data has emerged as a critical design challenge in the ultra-densely deployed cellular internet of things network (CITN) d…