1 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.CV2024★ 1 cited
Answering Diverse Questions via Text Attached with Key Audio-Visual Clues
Qilang Ye, Zitong Yu, Xin Liu
Audio-visual question answering (AVQA) requires reference to video content and auditory information, followed by correlating the question to predict the most precise answer. Althou…
cs.CV2024★ 1 cited
CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios
Qilang Ye, Zitong Yu, Rui Shao +3
This paper focuses on the challenge of answering questions in scenarios that are composed of rich and complex dynamic audio-visual components. Although existing Multimodal Large La…