14 papers
Imitation Learning for Robot Assistance in Open Surgery: A Multi-Policy Evaluation on Suture Following
Xucheng Wang, Zhizhou Yang, Xiaoman Zhang +3
This study presents the first evaluation of general-purpose imitation learning for surgeon-robot collaborative assistance in open surgery, targeting suture following: the grab-pull…
A systematic evaluation of vision-language models for observational astronomical reasoning tasks
Wenke Ren, Hengxiao Guo, Wenwen Zuo +1
Vision-language models (VLMs) are increasingly proposed as general-purpose tools for scientific data interpretation, yet their reliability on real astronomical observations across…
ReXSonoVQA: A Video QA Benchmark for Procedure-Centric Ultrasound Understanding
Xucheng Wang, Xiaoman Zhang, Sung Eun Kim +2
Ultrasound acquisition requires skilled probe manipulation and real-time adjustments. Vision-language models (VLMs) could enable autonomous ultrasound systems, but existing benchma…
ReXInTheWild: A Unified Benchmark for Medical Photograph Understanding
Oishi Banerjee, Sung Eun Kim, Alexandra N. Willauer +9
Everyday photographs taken with ordinary cameras are already widely used in telemedicine and other online health conversations, yet no comprehensive benchmark evaluates whether vis…
Do Mixed-Vendor Multi-Agent LLMs Improve Clinical Diagnosis?
Grace Chang Yuan, Xiaoman Zhang, Sung Eun Kim +1
Multi-agent large language model (LLM) systems have emerged as a promising approach for clinical diagnosis, leveraging collaboration among agents to refine medical reasoning. Howev…
ReX-MLE: The Autonomous Agent Benchmark for Medical Imaging Challenges
Roshan Kenia, Xiaoman Zhang, Pranav Rajpurkar
Autonomous coding agents built on large language models (LLMs) can now solve many general software and machine learning tasks, but they remain ineffective on complex, domain-specif…