5 papers
Token-Based Affordance Grounding with Large Vision-Language Models
Seung Il Lee, Qinqian Lei, Daguang Xu +4
Affordance grounding aims to localize image regions that support a specific action, serving as a core capability for physical intelligence and embodied perception. Previous studies…
Cosmos 3: Omnimodal World Models for Physical AI
NVIDIA, :, Aditi +293
We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-t…
AutoMedBench: Towards Medical AutoResearch with Agentic AI Models
Junqi Liu, Selena Song, Yuhan Wang +12
Autonomous agents are increasingly expected to support end-to-end medical-AI research workflows, moving beyond isolated prediction tasks or short-form clinical question answering.…
Surg: A Spectrum of Large-Scale Multimodal Data and Foundation Models for Surgical Intelligence
Zhitao Zeng, Mengya Xu, Jian Jiang +13
Surgical intelligence has the potential to improve the safety and consistency of surgical care, yet most existing surgical AI frameworks remain task-specific and struggle to genera…
LUMEN: Longitudinal Multi-Modal Radiology Model for Prognosis and Diagnosis
Zhifan Jiang, Dong Yang, Vishwesh Nath +7
Large vision-language models (VLMs) have evolved from general-purpose applications to specialized use cases such as in the clinical domain, demonstrating potential for decision sup…