1 paper
Morunliu Yang, Ruotao Xu, Le Li +6
Omnimodal large language models (OmniLLMs) have recently gained increasing attention for unified audio-video understanding. However, processing long multimodal token sequences intr…