11 papers
Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing
Junliang Ye, Kenkun Liu, Guocun Wang +13
Recent advances in image generation have demonstrated the potential of unified multimodal models that integrate understanding, generation, and editing. However, unified 3D modeling…
A Comparative Evaluation of Large Vision-Language Models for 2D Object Detection under SOTIF Conditions
Ji Zhou, Yilin Ding, Yongqi Zhao +4
The paper systematically evaluates ten large vision‑language models for 2D object detection on the PeSOTIF benchmark, comparing their recall and precision to YOLOv5 and RT‑DETRv4 u…
JointEdit3D: Feed-Forward 3D Scene Editing in a Unified Latent Space
Xinnan Zhu, Ruijie Xu, Jiayu Ying +4
Existing 3D scene editing methods typically rely on per-scene optimization over explicit 3D representations or cascaded edit-and-reconstruct pipelines, resulting in high test-time…
WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation
Baining Zhao, Jiacheng Xu, Weicheng Feng +13
Aerial vision-language navigation (VLN) requires agents to follow natural-language instructions through closed-loop perception and action in 3D environments. We argue that aerial V…
Forecast-aware Gaussian Splatting for Predictive 3D Representation in Language-Guided Pick-and-Place Manipulation
Kaixin Jia, Jiacheng Xu
We introduce Forecast-aware Gaussian Splatting (Forecast-GS), a predictive 3D representation framework for language-conditioned robotic manipulation. While recent manipulation syst…
How Far Are Large Multimodal Models from Human-Level Spatial Action? A Benchmark for Goal-Oriented Embodied Navigation in Urban Airspace
Baining Zhao, Ziyou Wang, Jianjie Fang +8
Large multimodal models (LMMs) show strong visual-linguistic reasoning but their capacity for spatial decision-making and action remains unclear. In this work, we investigate wheth…