13 papers
MentalThink: Shaping Thoughts in Mental SVG World
Kangheng Lin, Jisheng Yin, Dingming Li +11
We introduce MentalThink, a visual-symbolic reasoning paradigm that equips Multimodal LLMs (MLLMs) with an executable mechanism for "mental" visualization. The core of MentalThink…
PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception
Yana Wei, Hongbo Peng, Yanlin Lai +14
We introduce PerceptionRubrics, a rubric-based evaluation framework that addresses the gap between saturated benchmark scores and real-world brittleness. Shifting evaluation from h…
SpatialEvo: Self-Evolving Spatial Intelligence via Deterministic Geometric Environments
Dinging Li, Yingxiu Zhao, Xinrui Cheng +16
Spatial reasoning over three-dimensional scenes is a core capability for embodied intelligence, yet continuous model improvement remains bottlenecked by the cost of geometric annot…
Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters
Ailin Huang, Ang Li, Aobo Kong +213
We introduce Step 3.5 Flash, a sparse Mixture-of-Experts (MoE) model that bridges frontier-level agentic intelligence and computational efficiency. We focus on what matters most wh…
DM0: An Embodied-Native Vision-Language-Action Model towards Physical AI
En Yu, Haoran Lv, Jianjian Sun +46
Moving beyond the traditional paradigm of adapting internet-pretrained models to physical tasks, we present DM0, an Embodied-Native Vision-Language-Action (VLA) framework designed…
STEP3-VL-10B Technical Report
Ailin Huang, Chengyuan Yao, Chunrui Han +90
We present STEP3-VL-10B, a lightweight open-source foundation model designed to redefine the trade-off between compact efficiency and frontier-level multimodal intelligence. STEP3-…