1 paper · 1 filter
Zijian Fu, Changsheng Lv, Xianlin Zhang +2
In this paper, we propose a novel Multi-Modal Scene Graph with Kolmogorov-Arnold Expert Network for Audio-Visual Question Answering (SHRIKE). The task aims to mimic human reasoning…