2 citations · 4 across the 15 of their papers we have counts for
1 paper · 1 filter
Guangzhi Sun, Yixuan Li, Yudong Yang +1
Audio-visual large language models (LLMs) hold strong promise for long-form video understanding, yet their long-video inference is fundamentally limited by the linear growth of vid…