collaborators

6 papers

cs.CV2026

Cosmos 3: Omnimodal World Models for Physical AI

NVIDIA, :, Aditi +293

We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-t…

cs.AI2026

Surg: A Spectrum of Large-Scale Multimodal Data and Foundation Models for Surgical Intelligence

Zhitao Zeng, Mengya Xu, Jian Jiang +13

Surgical intelligence has the potential to improve the safety and consistency of surgical care, yet most existing surgical AI frameworks remain task-specific and struggle to genera…

cs.CV2026

Generalized Recognition of Basic Surgical Actions Enables Skill Assessment and Vision-Language-Model-based Surgical Planning

Mengya Xu, Daiyun Shen, Jie Zhang +19

Artificial intelligence, imaging, and large language models have the potential to transform surgical practice, training, and automation. Understanding and modeling of basic surgica…

cs.RO2025

Cosmos-H-Surgical: Learning Surgical Robot Policies from Videos via World Modeling

Yufan He, Pengfei Guo, Mengya Xu +11

Data scarcity remains a fundamental barrier to achieving fully autonomous surgical robots. While large scale vision language action (VLA) models have shown impressive generalizatio…

cs.CV2025

SAP-Bench: Benchmarking Multimodal Large Language Models in Surgical Action Planning

Mengya Xu, Zhongzhen Huang, Dillan Imans +3

Effective evaluation is critical for driving advancements in MLLM research. The surgical action planning (SAP) task, which aims to generate future action sequences from visual inpu…

cs.CL2025

Surgical Action Planning with Large Language Models

Mengya Xu, Zhongzhen Huang, Jie Zhang +2

In robot-assisted minimally invasive surgery, we introduce the Surgical Action Planning (SAP) task, which generates future action plans from visual inputs to address the absence of…