15 papers
AdaTurn: Budget-Aware Test-Time Scaling for Active Visual Perception Agents
Susan Liang, Chao Huang, Filippos Bellos +3
Active visual agents solve fine-grained image tasks by interleaving reasoning with image-grounding actions across multiple turns. However, deployment-time rollout budgets are rarel…
SurgAtlas: A Large-Scale Surgical Video-Language Dataset with 2,391 Hours of Open and Minimally Invasive Surgery
Filippos Bellos, Andre S. Gala-Garza, Miaowei Wang +8
We introduce SurgAtlas, the largest surgical video-language dataset to date, comprising 15,291 videos (2,391 hours) spanning 18 surgical specialties and over 5,000 procedure types,…
NTILC: Neural Tool Invocation via Learned Compression
Andrew Krikorian, Yayuan Li, Jason J. Corso
Agentic tool-calling language models depend on large registries of callable APIs, functions, and local actions. Placing full tool specifications directly in the prompt incurs a cos…
EchoVQA: Enabling Conversational Assistance for Point-of-Care Cardiac Ultrasound
Filippos Bellos, Yutong Li, Jessie N Dong +6
Point-of-care transthoracic echocardiography (TTE) enables cardiac assessment in virtually any clinical setting, yet its diagnostic utility remains constrained by the expertise req…
Substantial, Decomposable, and Invisible: Visual Context Misalignment in Instructional Videos for Physical Tasks
Yayuan Li, Chenglin Li, Jingying Wang +3
Instructional videos are the dominant medium for learning physical tasks, yet they rarely match the user's real-world visual context. Motor simulation and cognitive load theories p…
Follow Your Heart: Landmark-Guided Transducer Pose Scoring for Point-of-Care Echocardiography
Zaiyang Guo, Jessie N. Dong, Filippos Bellos +8
Point-of-care transthoracic echocardiography (TTE) makes it possible to assess a patient's cardiac function in almost any setting. A critical step in the TTE exam is acquisition of…