4 papers
Not All Modalities Are Equal: Instruction-Aware Gating for Multimodal Videos
Bonan Ding, Umair Nawaz, Ufaq Khan +5
Pre-trained video large language models excel at visual reasoning. However, they struggle when videos arrive with auxiliary streams, such as audio, depth map, or dense temporal evi…
WorldCache: Content-Aware Caching for Accelerated Video World Models
Umair Nawaz, Ahmed Heakl, Ufaq Khan +3
Diffusion Transformers (DiTs) power high-fidelity video world models but remain computationally expensive due to sequential denoising and costly spatio-temporal attention. Training…
Surgical Scene Understanding in the Era of Foundation AI Models: A Comprehensive Review
Ufaq Khan, Umair Nawaz, Adnan Qayyum +5
Recent advancements in machine learning (ML) and deep learning (DL), particularly through the introduction of Foundation Models (FMs), have significantly enhanced surgical scene un…
A Multimodal and Multi-centric Head and Neck Cancer Dataset for Segmentation, Diagnosis and Outcome Prediction
Numan Saeed, Salma Hassan, Shahad Hardan +40
We present a publicly available multimodal dataset for head and neck cancer research, comprising 1123 annotated Positron Emission Tomography/Computed Tomography (PET/CT) studies fr…