6 papers
Cosmos 3: Omnimodal World Models for Physical AI
NVIDIA, :, Aditi +293
We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-t…
Surg: A Spectrum of Large-Scale Multimodal Data and Foundation Models for Surgical Intelligence
Zhitao Zeng, Mengya Xu, Jian Jiang +13
Surgical intelligence has the potential to improve the safety and consistency of surgical care, yet most existing surgical AI frameworks remain task-specific and struggle to genera…
Generalized Recognition of Basic Surgical Actions Enables Skill Assessment and Vision-Language-Model-based Surgical Planning
Mengya Xu, Daiyun Shen, Jie Zhang +19
Artificial intelligence, imaging, and large language models have the potential to transform surgical practice, training, and automation. Understanding and modeling of basic surgica…
Cosmos-H-Surgical: Learning Surgical Robot Policies from Videos via World Modeling
Yufan He, Pengfei Guo, Mengya Xu +11
Data scarcity remains a fundamental barrier to achieving fully autonomous surgical robots. While large scale vision language action (VLA) models have shown impressive generalizatio…
SAP-Bench: Benchmarking Multimodal Large Language Models in Surgical Action Planning
Mengya Xu, Zhongzhen Huang, Dillan Imans +3
Effective evaluation is critical for driving advancements in MLLM research. The surgical action planning (SAP) task, which aims to generate future action sequences from visual inpu…
Surgical Action Planning with Large Language Models
Mengya Xu, Zhongzhen Huang, Jie Zhang +2
In robot-assisted minimally invasive surgery, we introduce the Surgical Action Planning (SAP) task, which generates future action plans from visual inputs to address the absence of…