6 papers
Text-Driven 3D Indoor Scene Synthesis in Non-Manhattan Environments
Xianhui Meng, Zirui Song, Yuchen Zhang +10
Large Language Models (LLMs) have demonstrated remarkable capabilities in 3D indoor synthesis for Manhattan environments. However, existing methods often fail to capture plausible…
Robust 3D Alignment of Generative Reconstructions via Partial Monocular Observations
Yuchen Zhang, Luanyuan Dai, Yiwei Wang +7
Aligning generative 3D reconstructions with partial monocular observations is a critical but under-explored challenge in computer vision. This task is inherently ill-posed due to s…
MiMo-Embodied: X-Embodied Foundation Model Technical Report
Xiaoshuai Hao, Lei Zhou, Zhijian Huang +41
We open-source MiMo-Embodied, the first cross-embodied foundation model to successfully integrate and achieve state-of-the-art performance in both Autonomous Driving and Embodied A…
DSBench: A Comprehensive Benchmark for Evaluating External and In-Cabin Risks
Xianhui Meng, Yuchen Zhang, Zhijian Huang +12
Vision-Language Models (VLMs) show great promise for autonomous driving, but their suitability for safety-critical scenarios is largely unexplored, raising safety concerns. This is…
Is your VLM Sky-Ready? A Comprehensive Spatial Intelligence Benchmark for UAV Navigation
Lingfeng Zhang, Yuchen Zhang, Hongsheng Li +7
Vision-Language Models (VLMs), leveraging their powerful visual perception and reasoning capabilities, have been widely applied in Unmanned Aerial Vehicle (UAV) tasks. However, the…
Team Xiaomi EV-AD VLA: Caption-Guided Retrieval System for Cross-Modal Drone Navigation -- Technical Report for IROS 2025 RoboSense Challenge Track 4
Lingfeng Zhang, Erjia Xiao, Yuchen Zhang +7
Cross-modal drone navigation remains a challenging task in robotics, requiring efficient retrieval of relevant images from large-scale databases based on natural language descripti…