4 papers
Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions
Saurav Sengupta, Nazanin Moradinasab, Jiebei Liu +1
Recent research suggests that Vision Language Models (VLMs) often rely on inherent biases learned during training when responding to queries about visual properties of images. Thes…
Combining Residual U-Net and Data Augmentation for Dense Temporal Segmentation of Spike Wave Discharges in Single-Channel EEG
Saurav Sengupta, Scott Kilianski, Suchetha Sharma +4
Manual annotation of spike-wave discharges (SWDs), the electrographic hallmark of absence seizures, is labor-intensive for long-term electroencephalography (EEG) monitoring studies…
Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features
Saurav Sengupta, Nazanin Moradinasab, Jiebei Liu +1
Recent research on Vision Language Models (VLMs) suggests that they rely on inherent biases learned during training to respond to questions about visual properties of an image. The…
Towards Robust Multimodal Representation: A Unified Approach with Adaptive Experts and Alignment
Nazanin Moradinasab, Saurav Sengupta, Jiebei Liu +2
Healthcare relies on multiple types of data, such as medical images, genetic information, and clinical records, to improve diagnosis and treatment. However, missing data is a commo…