5 citations · 12 across the 16 of their papers we have counts for
9 papers · 1 filter
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models
Sri Harsha Dumpala, David Arps, Sageev Oore +2
Vision-language models (VLMs), serve as foundation models for multi-modal applications such as image captioning and text-to-image generation. Recent studies have highlighted limita…
Sensitivity of Generative VLMs to Semantically and Lexically Altered Prompts
Sri Harsha Dumpala, Aman Jaiswal, Chandramouli Sastry +3
Despite the significant influx of prompt-tuning techniques for generative vision-language models (VLMs), it remains unclear how sensitive these models are to lexical and semantic a…
XANE Background Acoustic Embeddings: Ablation and Clustering Analysis
Dushyant Sharma, James Fosburgh, Sri Harsha Dumpala +3
We explore the recently proposed explainable acoustic neural embedding~(XANE) system that models the background acoustics of a speech signal in a non-intrusive manner. The XANE emb…
Self-Supervised Embeddings for Detecting Individual Symptoms of Depression
Sri Harsha Dumpala, Katerina Dikaios, Abraham Nunes +3
Depression, a prevalent mental health disorder impacting millions globally, demands reliable assessment systems. Unlike previous studies that focus solely on either detecting depre…
Predicting Individual Depression Symptoms from Acoustic Features During Speech
Sebastian Rodriguez, Sri Harsha Dumpala, Katerina Dikaios +3
Current automatic depression detection systems provide predictions directly without relying on the individual symptoms/items of depression as denoted in the clinical depression rat…
SUGARCREPE++ Dataset: Vision-Language Model Sensitivity to Semantic and Lexical Alterations
Sri Harsha Dumpala, Aman Jaiswal, Chandramouli Sastry +3
Despite their remarkable successes, state-of-the-art large language models (LLMs), including vision-and-language models (VLMs) and unimodal language models (ULMs), fail to understa…