collaborators

15 papers

cs.CV2026

SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models

Jing Wu, Jianhua Wu, Jiayi Guan +5

Vision-Language Models (VLMs) perform well on commonsense reasoning tasks but struggle with visual spatial reasoning. Most existing solutions introduce extra 3D prior inputs or ext…

cs.AI2026

Text-Driven 3D Indoor Scene Synthesis in Non-Manhattan Environments

Xianhui Meng, Zirui Song, Yuchen Zhang +10

Large Language Models (LLMs) have demonstrated remarkable capabilities in 3D indoor synthesis for Manhattan environments. However, existing methods often fail to capture plausible…

cs.RO2026

OneVLA: A Unified Framework for Embodied Tasks

Lingfeng Zhang, Xiaoshuai Hao, Yingbo Tang +10

Navigation and manipulation are fundamental capabilities of embodied intelligence, enabling robots to interpret natural language commands and interact physically with their surroun…

cs.LG2026

Reasoning emerges from constrained inference manifolds in large language models

Yanbiao Ma, Fei Luo, Linfeng Zhang +10

Reasoning in large language models is predominantly evaluated through labeled benchmarks, conflating task performance with the quality of internal inference. Here we study reasonin…

cs.AI2026

Thinking in Text and Images: Interleaved Vision--Language Reasoning Traces for Long-Horizon Robot Manipulation

Jinkun Liu, Haohan Chi, Lingfeng Zhang +6

Long-horizon robotic manipulation requires plans that are both logically coherent and geometrically grounded. Existing Vision-Language-Action policies usually hide planning in late…

cs.RO2026

MiMo-Embodied: X-Embodied Foundation Model Technical Report

Xiaoshuai Hao, Lei Zhou, Zhijian Huang +41

We open-source MiMo-Embodied, the first cross-embodied foundation model to successfully integrate and achieve state-of-the-art performance in both Autonomous Driving and Embodied A…