collaborators

10 papers

cs.CL2026

PragMatch: Separating Pragmatic Incongruity from Cross-Modal Mismatch in Large Vision-Language Models

Zhanna Mukhametsharip, Vera Demberg, Varsha Suresh

Large Vision-Language Models (LVLMs) have demonstrated strong performance on multimodal benchmarks, yet it remains unclear whether they genuinely reason about relationships between…

cs.CL2026

Semantic Motion Anchors: Bridging Motion and Meaning in Co-Speech Gestures

Varsha Suresh, Mohammad Mahdi Abootorabi, Mohamed Salman +5

Learning a shared representation between spoken text and gesture is central to co-speech gesture retrieval, synthesis, and understanding, but remains challenging for semantically m…

cs.AI2026

MuPHI: Learning Implicit Multimodal Harm Reasoning via Semantically Grounded Reward Optimization

Anisha Saha, Varsha Suresh, Teodora Kamova +3

Understanding how harm emerges from interaction between otherwise benign image-text pairs requires intent-aware cross-modal reasoning beyond surface-level features. Existing vision…

cs.CL2026

System-Mediated Attention Imbalances Make Vision-Language Models Say Yes

Tsan Tsai Chan, Varsha Suresh, Anisha Saha +2

Vision-language model (VLM) hallucination is commonly linked to imbalanced allocation of attention across input modalities: system, image and text. However, existing mitigation str…

cs.CL2026

Generation-Step-Aware Framework for Cross-Modal Representation and Control in Multilingual Speech-Text Models

Toshiki Nakai, Varsha Suresh, Vera Demberg

Multilingual speech-text models rely on cross-modal language alignment to transfer knowledge between speech and text, but it remains unclear whether this reflects shared computatio…

cs.CL2026

Modeling Turn-Taking with Semantically Informed Gestures

Varsha Suresh, M. Hamza Mughal, Christian Theobalt +1

In conversation, humans use multimodal cues, such as speech, gestures, and gaze, to manage turn-taking. While linguistic and acoustic features are informative, gestures provide com…