3 papers
cs.CL2026
Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis
Zicheng Kong, Dehua Ma, Zhenbo Xu +9
Multimodal large language models (MLLMs) struggle with alignment due to the limitations of existing reward models (RMs), which are predominantly vision-centric, dependent on costly…
cs.AI2025
OmniPlay: Benchmarking Omni-Modal Models on Omni-Modal Game Playing
Fuqing Bie, Shiyu Huang, Xijia Tao +6
While generalist foundation models like Gemini and GPT-4o demonstrate impressive multi-modal competence, existing evaluations fail to test their intelligence in dynamic, interactiv…
cs.LG2023
OpenRL: A Unified Reinforcement Learning Framework
Shiyu Huang, Wentse Chen, Yiwen Sun +2
We present OpenRL, an advanced reinforcement learning (RL) framework designed to accommodate a diverse array of tasks, from single-agent challenges to complex multi-agent systems.…