3 citations · 14 across the 21 of their papers we have counts for
4 papers · 1 filter
Show-Harness: Just a VLM Agent Can Play Robots
Yanzhe Chen, Zechen Bai, Zhijun Cao +7
Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence into robot control remains challenging. We present Show-Harne…
SurgNarrator: A Generative Retrieval Framework for Surgical Video Understanding
Yuqing Feng, Jiawei Ma, Kevin Qinghong Lin +6
Surgical procedures unfold as structured and recurring clinical events, whose real-time understanding via intraoperative surgical videos is critical for intraoperative decision-mak…
Demo2Tutorial: From Human Experience to Multimodal Software Tutorials
Zechen Bai, Zhiheng Chen, Yiqi Lin +5
Human experience in digital environments offers a vast, underexplored resource of authentic, untrimmed interactions that contain rich procedural knowledge. We introduce Demo2Tutori…
FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection
Mingyu Ouyang, Kevin Qinghong Lin, Mike Zheng Shou +1
Vision-Language Models (VLMs) have shown remarkable performance in User Interface (UI) grounding tasks, driven by their ability to process increasingly high-resolution screenshots.…