collaborators

7 papers

cs.CV2026

HistoSeg++: Delving deeper with attention and multiscale feature fusion for biomarker segmentation

Saad Wazir, Rao Faizan, Daeyoung Kim

Segmentation of biomarkers in medical images is frequently viewed as a first step towards medical image analysis in any bioinformatics or biomedical application. Despite progress,…

cs.CV2026

MedCAGD: Context-Aware Gated Decoder for Efficient Medical Image Segmentation

Saad Wazir, Patrick Dominique Vibild, Dinh Phu Tran +2

Medical image segmentation relies on the ability of encoder-decoder architectures to translate rich feature representations into accurate pixel-level predictions under challenging…

cs.CV2026

TB-AVA: Text as a Semantic Bridge for Audio-Visual Parameter Efficient Finetuning

Seongah Kim, Dinh Phu Tran, Hyeontaek Hwang +3

Audio-visual understanding requires effective alignment between heterogeneous modalities, yet cross-modal correspondence remains challenging when temporally aligned audio and visua…

cs.CV2026

SAT: Selective Aggregation Transformer for Image Super-Resolution

Dinh Phu Tran, Thao Do, Saad Wazir +3

Transformer-based approaches have revolutionized image super-resolution by modeling long-range dependencies. However, the quadratic computational complexity of vanilla self-attenti…

cs.LG2026

Knowing When to Answer: Adaptive Confidence Refinement for Reliable Audio-Visual Question Answering

Dinh Phu Tran, Jihoon Jeong, Saad Wazir +4

We present a formal problem formulation for \textit{Reliable} Audio-Visual Question Answering (-AVQA), where we prefer abstention over answering incorrectly. While rec…

eess.IV2025

Rethinking Decoder Design: Improving Biomarker Segmentation Using Depth-to-Space Restoration and Residual Linear Attention

Saad Wazir, Daeyoung Kim

Segmenting biomarkers in medical images is crucial for various biotech applications. Despite advances, Transformer and CNN based methods often struggle with variations in staining…