12 citations · 15 across the 13 of their papers we have counts for
1 paper · 1 filter
Guangzheng Hu, Ziyue Jiang, Weixu Qiao +14
Multimodal understanding models that can jointly judge text-to-image (T2I), text-to-video (T2V) and text-to-speech (TTS) generation are increasingly used as "OmniJudges" for evalua…