2 citations · 2 across the 3 of their papers we have counts for
3 papers
cs.SD2025
FSSUAVL: A Discriminative Framework using Vision Models for Federated Self-Supervised Audio and Image Understanding
Yasar Abbas Ur Rehman, Kin Wai Lau, Yuyang Xie +2
Recent studies have demonstrated that vision models can effectively learn multimodal audio-image representations when paired. However, the challenge of enabling deep models to lear…
eess.IV2024
Acquire Precise and Comparable Fundus Image Quality Score: FTHNet and FQS Dataset
Zheng Gong, Zhuo Deng, Run Gan +8
The retinal fundus images are utilized extensively in the diagnosis, and their quality can directly affect the diagnosis results. However, due to the insufficient dataset and algor…
cs.SD2023★ 2 cited
AudioInceptionNeXt: TCL AI LAB Submission to EPIC-SOUND Audio-Based-Interaction-Recognition Challenge 2023
Kin Wai Lau, Yasar Abbas Ur Rehman, Yuyang Xie +1
This report presents the technical details of our submission to the 2023 Epic-Kitchen EPIC-SOUNDS Audio-Based Interaction Recognition Challenge. The task is to learn the mapping fr…