12 papers
MentalThink: Shaping Thoughts in Mental SVG World
Kangheng Lin, Jisheng Yin, Dingming Li +11
We introduce MentalThink, a visual-symbolic reasoning paradigm that equips Multimodal LLMs (MLLMs) with an executable mechanism for "mental" visualization. The core of MentalThink…
PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception
Yana Wei, Hongbo Peng, Yanlin Lai +14
We introduce PerceptionRubrics, a rubric-based evaluation framework that addresses the gap between saturated benchmark scores and real-world brittleness. Shifting evaluation from h…
STEP3-VL-10B Technical Report
Ailin Huang, Chengyuan Yao, Chunrui Han +90
We present STEP3-VL-10B, a lightweight open-source foundation model designed to redefine the trade-off between compact efficiency and frontier-level multimodal intelligence. STEP3-…
Incremental Human-Object Interaction Detection with Invariant Relation Representation Learning
Yana Wei, Zeen Chi, Chongyu Wang +4
In open-world environments, human-object interactions (HOIs) evolve continuously, challenging conventional closed-world HOI detection models. Inspired by humans' ability to progres…
World-in-World: World Models in a Closed-Loop World
Jiahan Zhang, Muqing Jiang, Nanru Dai +14
Generative world models (WMs) can now simulate worlds with striking visual realism, which naturally raises the question of whether they can endow embodied agents with predictive pe…
Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
Yana Wei, Liang Zhao, Jianjian Sun +15
The remarkable reasoning capability of large language models (LLMs) stems from cognitive behaviors that emerge through reinforcement with verifiable rewards. This work investigates…