collaborators

5 papers

cs.AI2026

Vision Language Models Cannot Reason About Physical Transformation

Dezhi Luo, Yijiang Li, Maijunxian Wang +7

Understanding physical transformations is fundamental for reasoning in dynamic environments. While Vision Language Models (VLMs) show promise in embodied applications, whether they…

cs.CL2026

CocoaBench: Evaluating Unified Digital Agents in the Wild

CocoaBench Team, Shibo Hao, Zhining Zhang +29

LLM agents now perform strongly in software engineering, deep research, GUI automation, and various other applications, while recent agent scaffolds and models are increasingly int…

cs.AI2026

SimWorld-Robotics: Synthesizing Photorealistic and Dynamic Urban Environments for Multimodal Robot Navigation and Collaboration

Yan Zhuang, Jiawei Ren, Xiaokang Ye +9

Recent advances in foundation models have shown promising results in developing generalist robotics that can perform diverse tasks in open-ended scenarios given multimodal inputs.…

cs.AI2026

SimWorld: An Open-ended Realistic Simulator for Autonomous Agents in Physical and Social Worlds

Jiawei Ren, Yan Zhuang, Xiaokang Ye +20

While LLM/VLM-powered AI agents have advanced rapidly in math, coding, and computer use, their applications in complex physical and social environments remain challenging. Building…

cs.AI2025

DeliveryBench: Can Agents Earn Profit in Real World?

Lingjun Mao, Jiawei Ren, Kun Zhou +3

LLMs and VLMs are increasingly deployed as embodied agents, yet existing benchmarks largely revolve around simple short-term tasks and struggle to capture rich realistic constraint…