3 papers
cs.CL2025
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution
Jiahui Li, Lin Li, Tai-wei Chang +4
Reinforcement learning from human feedback (RLHF) offers a promising approach to aligning large language models (LLMs) with human preferences. Typically, a reward model is trained…
cs.LG2025
M2-omni: Advancing Omni-MLLM for Comprehensive Modality Support with Competitive Performance
Qingpei Guo, Kaiyou Song, Zipeng Feng +9
We present M2-omni, a cutting-edge, open-source omni-MLLM that achieves competitive performance to GPT-4o. M2-omni employs a unified multimodal sequence modeling framework, which e…
cs.LG2025
Learning Causal Transition Matrix for Instance-dependent Label Noise
Jiahui Li, Tai-Wei Chang, Kun Kuang +3
Noisy labels are both inevitable and problematic in machine learning methods, as they negatively impact models' generalization ability by causing overfitting. In the context of lea…