3 papers
cs.CV2026
EgoMAGIC- An Egocentric Video Field Medicine Dataset for Training Perception Algorithms
Brian VanVoorst, Nicholas Walczak, Christopher Gilleo +9
This paper introduces EgoMAGIC (Medical Assistance, Guidance, Instruction, and Correction), an egocentric medical activity dataset collected as part of DARPA's Perceptually-enabled…
cs.CV2026
MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding
Yuhao Su, Anwesa Choudhuri, Zhongpai Gao +8
Large vision-language models struggle with medical video understanding, where spatial precision, temporal reasoning, and clinical semantics are critical. To address this, we first…
cs.CV2025
DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long Videos
Zijia Lu, A S M Iftekhar, Gaurav Mittal +6
Long Video Temporal Grounding (LVTG) aims at identifying specific moments within lengthy videos based on user-provided text queries for effective content retrieval. The approach ta…