4 papers
Beyond Referring Expressions: Scenario Comprehension Visual Grounding
Ruozhen He, Nisarg A. Shah, Qihua Dong +3
Existing visual grounding benchmarks primarily evaluate alignment between image regions and literal referring expressions, where models can often succeed by matching a prominent na…
Cinéaste: A Fine-grained Contextual Movie Question Answering Benchmark
Nisarg A. Shah, Amir Ziai, Chaitanya Ekanadham +1
While recent advancements in vision-language models have improved video understanding, diagnosing their capacity for deep, narrative comprehension remains a challenge. Existing ben…
StepAL: Step-aware Active Learning for Cataract Surgical Videos
Nisarg A. Shah, Bardia Safaei, Shameema Sikder +2
Active learning (AL) can reduce annotation costs in surgical video analysis while maintaining model performance. However, traditional AL methods, developed for images or short vide…
:~Cataract Surgical Masked Autoencoder (MAE) based Pre-training
Nisarg A. Shah, Wele Gedara Chaminda Bandara, Shameema Skider +2
Automated analysis of surgical videos is crucial for improving surgical training, workflow optimization, and postoperative assessment. We introduce a CSMAE, Masked Autoencoder (MAE…