1 paper
Guangzheng Hu, Ziyue Jiang, Weixu Qiao +14
Multimodal understanding models that can jointly judge text-to-image (T2I), text-to-video (T2V) and text-to-speech (TTS) generation are increasingly used as "OmniJudges" for evalua…