2 citations · 3 across the 11 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs
Chao Gong, Depeng Wang, Zhipeng Wei +3
Audio-Visual Large Language Models (AV-LLMs) face prohibitive computational costs of processing massive, redundant audio-visual tokens. Existing unimodal compression techniques fai…
cs.CV2023★ 2 cited
LayoutMask: Enhance Text-Layout Interaction in Multi-modal Pre-training for Document Understanding
Yi Tu, Ya Guo, Huan Chen +1
Visually-rich Document Understanding (VrDU) has attracted much research attention over the past years. Pre-trained models on a large number of document images with transformer-base…