10 papers
PragMatch: Separating Pragmatic Incongruity from Cross-Modal Mismatch in Large Vision-Language Models
Zhanna Mukhametsharip, Vera Demberg, Varsha Suresh
Large Vision-Language Models (LVLMs) have demonstrated strong performance on multimodal benchmarks, yet it remains unclear whether they genuinely reason about relationships between…
Semantic Motion Anchors: Bridging Motion and Meaning in Co-Speech Gestures
Varsha Suresh, Mohammad Mahdi Abootorabi, Mohamed Salman +5
Learning a shared representation between spoken text and gesture is central to co-speech gesture retrieval, synthesis, and understanding, but remains challenging for semantically m…
MuPHI: Learning Implicit Multimodal Harm Reasoning via Semantically Grounded Reward Optimization
Anisha Saha, Varsha Suresh, Teodora Kamova +3
Understanding how harm emerges from interaction between otherwise benign image-text pairs requires intent-aware cross-modal reasoning beyond surface-level features. Existing vision…
System-Mediated Attention Imbalances Make Vision-Language Models Say Yes
Tsan Tsai Chan, Varsha Suresh, Anisha Saha +2
Vision-language model (VLM) hallucination is commonly linked to imbalanced allocation of attention across input modalities: system, image and text. However, existing mitigation str…
Generation-Step-Aware Framework for Cross-Modal Representation and Control in Multilingual Speech-Text Models
Toshiki Nakai, Varsha Suresh, Vera Demberg
Multilingual speech-text models rely on cross-modal language alignment to transfer knowledge between speech and text, but it remains unclear whether this reflects shared computatio…
Modeling Turn-Taking with Semantically Informed Gestures
Varsha Suresh, M. Hamza Mughal, Christian Theobalt +1
In conversation, humans use multimodal cues, such as speech, gestures, and gaze, to manage turn-taking. While linguistic and acoustic features are informative, gestures provide com…