2 papers
cs.RO2026
MachEmbodied-U0: Unified Understanding and Generation Model for Embodied Intelligence
Haoran Wen, Wenfu Wang, Kunsong Shi +15
General-purpose robot control requires models to understand task intent, identify where to interact, capture how the scene evolves, and generate precise actions. Vision-language-ac…
cs.CV2026
LiAuto-MindViT: A Hybrid Vision Backbone with Adaptive Bidirectional Mamba
Lifu Mu, Shuai Chen, Wen Zheng +7
While Mamba-based models have shown strong potential for long sequence modeling, adapting them to vision is challenging due to the requirement of local neighborhood correlations an…