2 papers
cs.AI2026
Towards Generalizable Visually Grounded Exploration of Household Devices
Linhao Zheng, Zeming Liu, Wangke Chen +4
Recent advancements in Vision-Language Models (VLMs) have demonstrated impressive capabilities in static visual recognition and high-level semantic reasoning. However, current embo…
cs.RO2026
HODAgent: Towards On-Demand, Responsive Humanoids for Physical World Human Interaction
Wang Warren Chen, Jiahao Zhang, Zhenjiang Li +6
We propose HODAgent, a System-2 embodied agent for humanoid robots in service settings, addressing situated intent, responsive execution, task revision, and outcome verification. I…