1 paper
Yifan Shen, Pei Tian, Xinzhuo Li +8
Omni-modal models can ingest video, audio, and text, but unified access to multiple modalities does not guarantee that a model uses the right evidence. This gap is especially prono…