1 paper
Fuqing Bie, Shiyu Huang, Xijia Tao +6
While generalist foundation models like Gemini and GPT-4o demonstrate impressive multi-modal competence, existing evaluations fail to test their intelligence in dynamic, interactiv…