1 paper · 1 filter
Andrew Rouditchenko, Saurabhchand Bhati, Edson Araujo +4
We propose Omni-R1 which fine-tunes a recent multi-modal LLM, Qwen2.5-Omni, on an audio question answering dataset with the reinforcement learning method GRPO. This leads to new St…