6 citations · 14 across the 8 of their papers we have counts for
1 paper · 1 filter
Guangzhi Sun, Wenyi Yu, Changli Tang +7
Speech understanding as an element of the more generic video understanding using audio-visual large language models (av-LLMs) is a crucial yet understudied aspect. This paper propo…