1 paper · 1 filter
Chao Gong, Depeng Wang, Zhipeng Wei +3
Audio-Visual Large Language Models (AV-LLMs) face prohibitive computational costs of processing massive, redundant audio-visual tokens. Existing unimodal compression techniques fai…