4 papers
Latent World Models with Monotone Planning Costs for Image-Goal Navigation
Amirhosein Chahe, Siwei Cai, Lifeng Zhou
Image-goal navigation with latent world models requires not only accurate future prediction, but also a planning cost that reliably ranks candidate action sequences. We define the…
Threading Optimization for Vision-Language-Action Model Inference in Low-Cost Smart Agricultural Manipulation
Keith Truongcao, Christopher Nhu, Zijian An +3
Vision-Language Action (VLA) models continue to face challenges such as slow inference speed and difficulty performing fine-grained motion adjustments, limiting their widespread ad…
GA3T: A Ground-Aerial Terrain Traversability Dataset for Heterogeneous Robot Teams in Unstructured Environments
Siwei Cai, Knut Peterson, Quan Tran +7
Heterogeneous air-ground robot teams combine complementary sensing modalities, mobility characteristics, and spatial viewpoints that can significantly enhance perception in complex…
LLM-Land: Large Language Models for Context-Aware Drone Landing
Siwei Cai, Yuwei Wu, Lifeng Zhou
Autonomous landing is essential for drones deployed in emergency deliveries, post-disaster response, and other large-scale missions. By enabling self-docking on charging platforms,…