2 citations · 2 across the 4 of their papers we have counts for
4 papers
PRACTICE: From Experience to Expertise in Self-Evolving Embodied Agents
Ziyi Bai, Siqi Li, Tinglei Huang +1
Recent studies have shown that multimodal large language models (MLLMs) can serve as embodied agents, translating language instructions and visual observations into executable plan…
PRISM: : Planning and Reasoning with Intent in Simulated Embodied Environments
Yunn Kang Lim, Pengzhan Sun, Ziyi Bai +4
When an LLM-based embodied agent fails at a household task, the culprit could be misidentified objects, forgotten sub-goals, or poor action sequencing -- yet existing benchmarks re…
EgoActor: Grounding Task Planning into Spatial-aware Egocentric Actions for Humanoid Robots via Visual-Language Models
Yu Bai, MingMing Yu, Chaojie Li +3
Deploying humanoid robots in real-world settings is fundamentally challenging, as it demands tight integration of perception, locomotion, and manipulation under partial-information…
Glance and Focus: Memory Prompting for Multi-Event Video Question Answering
Ziyi Bai, Ruiping Wang, Xilin Chen
Video Question Answering (VideoQA) has emerged as a vital tool to evaluate agents' ability to understand human daily behaviors. Despite the recent success of large vision language…