1 citations · 1 across the 1 of their papers we have counts for
1 paper
Qilang Ye, Zitong Yu, Rui Shao +3
This paper focuses on the challenge of answering questions in scenarios that are composed of rich and complex dynamic audio-visual components. Although existing Multimodal Large La…