7 citations · 10 across the 7 of their papers we have counts for
1 paper · 1 filter
Ankan Deria, Hanoona Rasheed, Xilin He +2
Audio-visual understanding remains challenging because models must jointly interpret spoken content, visual events, and their temporal relationships. Existing omni models typically…