activity
20242026
collaborators
Showing 2026Show all

6 papers · 1 filter

cs.RO2026

Geometry Conditioning in an Embodied SLM: Training Controls and Robustness Diagnostics in a 0.8B Hybrid Model

Hao Li, Haofei Sun, Lin He

We study how physical-state inputs affect a 0.8B hybrid language model adapted for manipulation with 6.2M trainable parameters. Six conditions are trained on three LIBERO-Spatial t…

cs.CV2026

GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding

Hao Li, Han Fang, Zixin Pan +8

Although multimodal large language models (MLLMs) have achieved remarkable progress, understanding 3D spatial relationships from 2D images remains a critical challenge. Existing me…

cs.CV2026

Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding

Tianyi Gao, Han Fang, Tianyi Ding +9

Multimodal Large Language Models have made great progress in grounding tasks, yet existing methods still struggle to unify precise localization and complex reasoning. For one thing…

cs.CV2026

DrivingDepth: Sparse-Prompted Pixel-wise Scale Correction for Driving Depth Estimation

Chi Huang, Wenhao Zhang, Hang Yin +5

Dense depth estimation for autonomous driving faces a geometry-scale conflict: depth foundation models deliver pixel-aligned dense visual geometry without reliable metric scale, wh…

cs.CV2026

Monte Carlo Energy Aggregation for Mobile 3D Gaussian Splatting

Xiaobiao Du, YuAn Wang, Hao Li +3

Recent advances in 3D Gaussian Splatting have demonstrated unprecedented success in novel view synthesis. However, the substantial inference and storage overhead driven by high-ord…

cs.RO2026

Wow, wo, val! A Comprehensive Embodied World Model Evaluation Turing Test

Chun-Kai Fan, Xiaowei Chi, Xiaozhu Ju +18

As world models gain momentum in Embodied AI, an increasing number of works explore using video foundation models as predictive world models for downstream embodied tasks like 3D p…