4 papers · 1 filter
Multi-view Video-Pose Pretraining for Operating Room Surgical Activity Recognition
Idris Hamoud, Vinkle Srivastav, Muhammad Abdullah Jamal +3
Understanding the workflow of surgical procedures in complex operating rooms requires a deep understanding of the interactions between clinicians and their environment. Surgical ac…
A Two-Stage Progressive Pre-training using Multi-Modal Contrastive Masked Autoencoders
Muhammad Abdullah Jamal, Omid Mohareri
In this paper, we propose a new progressive pre-training method for image understanding tasks which leverages RGB-D datasets. The method utilizes Multi-Modal Contrastive Masked Aut…
VidLPRO: A eo-anguage re-training Framework for botic and Laparoscopic Surgery
Mohammadmahdi Honarmand, Muhammad Abdullah Jamal, Omid Mohareri
We introduce VidLPRO, a novel video-language (VL) pre-training framework designed specifically for robotic and laparoscopic surgery. While existing surgical VL models primarily rel…
Rethinking Class-Balanced Methods for Long-Tailed Visual Recognition from a Domain Adaptation Perspective
Muhammad Abdullah Jamal, Matthew Brown, Ming-Hsuan Yang +2
Object frequency in the real world often follows a power law, leading to a mismatch between datasets with long-tailed class distributions seen by a machine learning model and our e…