1 citations · 2 across the 4 of their papers we have counts for
4 papers
MA-AVT: Modality Alignment for Parameter-Efficient Audio-Visual Transformers
Tanvir Mahmud, Shentong Mo, Yapeng Tian +1
Recent advances in pre-trained vision transformers have shown promise in parameter-efficient audio-visual learning without audio pre-training. However, few studies have investigate…
Weakly-supervised Audio Separation via Bi-modal Semantic Similarity
Tanvir Mahmud, Saeed Amizadeh, Kazuhito Koishida +1
Conditional sound separation in multi-source audio mixtures without having access to single source sound data during training is a long standing challenge. Existing mix-and-separat…
SSVOD: Semi-Supervised Video Object Detection with Sparse Annotations
Tanvir Mahmud, Chun-Hao Liu, Burhaneddin Yaman +1
Despite significant progress in semi-supervised learning for image object detection, several key issues are yet to be addressed for video object detection: (1) Achieving good perfo…
CIFF-Net: Contextual Image Feature Fusion for Melanoma Diagnosis
Md Awsafur Rahman, Bishmoy Paul, Tanvir Mahmud +1
Melanoma is considered to be the deadliest variant of skin cancer causing around 75\% of total skin cancer deaths. To diagnose Melanoma, clinicians assess and compare multiple skin…