7 papers
HistoSeg++: Delving deeper with attention and multiscale feature fusion for biomarker segmentation
Saad Wazir, Rao Faizan, Daeyoung Kim
Segmentation of biomarkers in medical images is frequently viewed as a first step towards medical image analysis in any bioinformatics or biomedical application. Despite progress,…
MedCAGD: Context-Aware Gated Decoder for Efficient Medical Image Segmentation
Saad Wazir, Patrick Dominique Vibild, Dinh Phu Tran +2
Medical image segmentation relies on the ability of encoder-decoder architectures to translate rich feature representations into accurate pixel-level predictions under challenging…
TB-AVA: Text as a Semantic Bridge for Audio-Visual Parameter Efficient Finetuning
Seongah Kim, Dinh Phu Tran, Hyeontaek Hwang +3
Audio-visual understanding requires effective alignment between heterogeneous modalities, yet cross-modal correspondence remains challenging when temporally aligned audio and visua…
SAT: Selective Aggregation Transformer for Image Super-Resolution
Dinh Phu Tran, Thao Do, Saad Wazir +3
Transformer-based approaches have revolutionized image super-resolution by modeling long-range dependencies. However, the quadratic computational complexity of vanilla self-attenti…
Knowing When to Answer: Adaptive Confidence Refinement for Reliable Audio-Visual Question Answering
Dinh Phu Tran, Jihoon Jeong, Saad Wazir +4
We present a formal problem formulation for \textit{Reliable} Audio-Visual Question Answering (-AVQA), where we prefer abstention over answering incorrectly. While rec…
Rethinking Decoder Design: Improving Biomarker Segmentation Using Depth-to-Space Restoration and Residual Linear Attention
Saad Wazir, Daeyoung Kim
Segmenting biomarkers in medical images is crucial for various biotech applications. Despite advances, Transformer and CNN based methods often struggle with variations in staining…