74 citations · 549 across the 47 of their papers we have counts for
1 paper · 1 filter
Ziyi Yang, Mahmoud Khademi, Yichong Xu +16
The convergence of text, visual, and audio data is a key step towards human-like artificial intelligence, however the current Vision-Language-Speech landscape is dominated by encod…