4 papers · 1 filter
Perception-R1: Pioneering Perception Policy with Reinforcement Learning
En Yu, Kangheng Lin, Liang Zhao +11
Inspired by the success of DeepSeek-R1, we explore the potential of rule-based reinforcement learning (RL) in MLLM post-training for perception policy learning. While promising, ou…
Unhackable Temporal Rewarding for Scalable Video MLLMs
En Yu, Kangheng Lin, Liang Zhao +8
In the pursuit of superior video-processing MLLMs, we have encountered a perplexing paradox: the "anti-scaling law", where more data and larger models lead to worse performance. Th…
MEGL: Multimodal Explanation-Guided Learning
Yifei Zhang, Tianxu Jiang, Bo Pan +3
Explaining the decision-making processes of Artificial Intelligence (AI) models is crucial for addressing their "black box" nature, particularly in tasks like image classification.…
MLS-Track: Multilevel Semantic Interaction in RMOT
Zeliang Ma, Song Yang, Zhe Cui +4
The new trend in multi-object tracking task is to track objects of interest using natural language. However, the scarcity of paired prompt-instance data hinders its progress. To ad…