1 paper
Yasar Abbas Ur Rehman, Kin Wai Lau, Yuyang Xie +2
Recent studies have demonstrated that vision models can effectively learn multimodal audio-image representations when paired. However, the challenge of enabling deep models to lear…