11 papers
SafeRelBench: A Spatial-Relation-Aware Benchmark for Process-Level Safety in VLM-Driven Embodied Agents
Huaigang Yang, Ya Li, Min Ren +3
Vision-language models (VLMs) are increasingly used as the reasoning backbone of embodied agents, enabling robots to interpret visual scenes, follow language instructions, and plan…
BotDirector: Robot Storytelling Across the Symmetrical Reality with Multi-modal Interactions
Zhe Sun, Meng Wang, Lei Wang +4
Robot storytelling offers a unique blend of technological innovation and creative expression that engages children in unprecedented ways. However, the technical aspects are often t…
NormAct: Benchmarking Embodied Agents' Proactive Compliance with Unspoken Social Norms
Shiyun Zhao, Xinwei Song, Tianyu Guo +7
Embodied agents driven by multimodal large language models (MLLMs) can often complete everyday tasks from visual observations, but goal achievement does not establish whether they…
TeachAnything: A Multimodal Crowdsourcing Platform for Training Embodied AI Agents in Symmetrical Reality
Zidong Liu, Rongkai Liu, Yue Li +1
Symmetrical Reality (SR) is emerging as a future trend for human-agent coexistence, placing higher demands on agents to acquire human-like intelligence. It calls for richer and mor…
Exploring Human-Machine Coexistence in Symmetrical Reality
Zhenliang Zhang
In the context of the evolution of artificial intelligence (AI), the interaction between humans and AI entities has become increasingly salient, challenging the conventional human-…
TongSIM: A General Platform for Simulating Intelligent Machines
Zhe Sun, Kunlun Wu, Chuanjian Fu +24
As artificial intelligence (AI) rapidly advances, especially in multimodal large language models (MLLMs), research focus is shifting from single-modality text processing to the mor…