Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs
Chao Gong, Depeng Wang, Zhipeng Wei +3
Audio-Visual Large Language Models (AV-LLMs) face prohibitive computational costs of processing massive, redundant audio-visual tokens. Existing unimodal compression techniques fai…
cs.CV2024
UNK-VQA: A Dataset and a Probe into the Abstention Ability of Multi-modal Large Models
Yangyang Guo, Fangkai Jiao, Zhiqi Shen +2
Teaching Visual Question Answering (VQA) models to refrain from answering unanswerable questions is necessary for building a trustworthy AI system. Existing studies, though have ex…