1 paper · 1 filter
Yifan Zhu, Xinyu Mu, Tao Feng +3
Long-horizon omnimodal question answering answers questions by reasoning over text, images, audio, and video. Despite recent progress on OmniLLMs, low-resource long audio-video QA…