cloud-edge computing 1gesture recognition 1human-robot interaction 1large language models 1multimodal fusion 1vision-language models 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.RO2026
G0.5: One Autoregressive Stream for Robot Reasoning and Action
Yicheng Liu, Zibin Dong, Baijun Ye +24
The prevailing recipe for Vision-Language-Action (VLA) models couples a pretrained VLM with a separately trained flow-matching action expert. This makes the VLM a context encoder r…
cs.RO2026
An Intelligent-Cloud Edge Multimodal Interaction System for Robots
Zihan Guo, Xiaoqi Li
The paper proposes a cloud‑edge framework that combines an enhanced YOLO‑based gesture detector with coordinated large language model and vision‑language model agents to enable rob…