1 paper
Zhaolu Kang, Meixin Wu, Yu Xue +6
Omni-modal large language models are increasingly evaluated on clean text--vision--audio inputs, where every channel is present, synchronized, and readily interpretable. Such score…