1 paper
Jianghan Chao, Jianzhang Gao, Wenhui Tan +3
Understanding videos inherently requires reasoning over both visual and auditory information. To properly evaluate Omni-Large Language Models (Omni-LLMs), which are capable of proc…