collaborators

7 papers

cs.AI2026

ENVS: Environment-Native Verified Search for Long-Horizon GUI Agents

Yincheng Zhou, Athena Zhuoming Zhong, Shijie Zhang +3

As multimodal agents move from interface understanding to real software control, successful trajectory discovery in live desktop environments becomes a key challenge. GUI tasks req…

cs.RO2026

SafeDojo: Safe Reinforcement Learning for VLA via Interactive World Model

Kai Tang, Peidong Jia, Zhong Chu +15

Safe control is a prerequisite for real-world embodied intelligence, for which safe reinforcement learning has emerged as a promising paradigm. However, existing safe reinforcement…

cs.LG2026

Conf-Gen: Conformal Uncertainty Quantification for Generative Models

Gabriel Loaiza-Ganem, Kevin Zhang, Wei Cui +2

Conformal prediction (CP) and its extension, conformal risk control (CRC), are established frameworks for quantifying uncertainty in supervised machine learning through formal guar…

cs.CV2026

PhysicsMind: Sim and Real Mechanics Benchmarking for Physical Reasoning and Prediction in Foundational VLMs and World Models

Chak-Wing Mak, Guanyu Zhu, Boyi Zhang +16

Modern foundational Multimodal Large Language Models (MLLMs) and video world models have advanced significantly in mathematical, common-sense, and visual reasoning, but their grasp…

cs.RO2026

Wow, wo, val! A Comprehensive Embodied World Model Evaluation Turing Test

Chun-Kai Fan, Xiaowei Chi, Xiaozhu Ju +18

As world models gain momentum in Embodied AI, an increasing number of works explore using video foundation models as predictive world models for downstream embodied tasks like 3D p…

cs.RO2025

WoW: Towards a World omniscient World model Through Embodied Interaction

Xiaowei Chi, Peidong Jia, Chun-Kai Fan +33

Humans develop an understanding of intuitive physics through active interaction with the world. This approach is in stark contrast to current video models, such as Sora, which rely…