1 paper
Seunghee Kim, Bumkyu Park, Kyudan Jung +5
Most testbeds for omni-modal models assess multimodal understanding via textual outputs, leaving it unclear whether these models can properly speak their answers. To study this, we…