2 papers
cs.CV2026
VisGym: Diverse, Customizable, Scalable Environments for Multimodal Agents
Zirui Wang, Junyi Zhang, Jiaxin Ge +9
Modern Vision-Language Models (VLMs) remain poorly characterized in multi-step visual interactions, particularly in how they integrate perception, memory, and action over long hori…
cs.RO2024
FogROS2-FT: Fault Tolerant Cloud Robotics
Kaiyuan Chen, Kush Hari, Trinity Chung +8
Cloud robotics enables robots to offload complex computational tasks to cloud servers for performance and ease of management. However, cloud compute can be costly, cloud services c…