1k citations · 1.3k across the 10 of their papers we have counts for
1 paper · 1 filter
Ziyi Yang, Mahmoud Khademi, Yichong Xu +16
The convergence of text, visual, and audio data is a key step towards human-like artificial intelligence, however the current Vision-Language-Speech landscape is dominated by encod…