1 paper
Bhavesh Kumar, Dylan Feng, Leonard Tang
Multimodal judges struggle to ground decisions in visual evidence. We present MJ1, a multimodal judge trained with reinforcement learning that enforces visual grounding through a s…