activity
20242026
collaborators

10 papers

cs.CL2026

Semantic Motion Anchors: Bridging Motion and Meaning in Co-Speech Gestures

Varsha Suresh, Mohammad Mahdi Abootorabi, Mohamed Salman +5

Learning a shared representation between spoken text and gesture is central to co-speech gesture retrieval, synthesis, and understanding, but remains challenging for semantically m…

cs.AI2026

MuPHI: Learning Implicit Multimodal Harm Reasoning via Semantically Grounded Reward Optimization

Anisha Saha, Varsha Suresh, Teodora Kamova +3

Understanding how harm emerges from interaction between otherwise benign image-text pairs requires intent-aware cross-modal reasoning beyond surface-level features. Existing vision…

cs.CL2026

System-Mediated Attention Imbalances Make Vision-Language Models Say Yes

Tsan Tsai Chan, Varsha Suresh, Anisha Saha +2

Vision-language model (VLM) hallucination is commonly linked to imbalanced allocation of attention across input modalities: system, image and text. However, existing mitigation str…

cs.CL2026

Generation-Step-Aware Framework for Cross-Modal Representation and Control in Multilingual Speech-Text Models

Toshiki Nakai, Varsha Suresh, Vera Demberg

Multilingual speech-text models rely on cross-modal language alignment to transfer knowledge between speech and text, but it remains unclear whether this reflects shared computatio…

cs.CL2026

Modeling Turn-Taking with Semantically Informed Gestures

Varsha Suresh, M. Hamza Mughal, Christian Theobalt +1

In conversation, humans use multimodal cues, such as speech, gestures, and gaze, to manage turn-taking. While linguistic and acoustic features are informative, gestures provide com…

cs.LG2025

MUStReason: A Benchmark for Diagnosing Pragmatic Reasoning in Video-LMs for Multimodal Sarcasm Detection

Anisha Saha, Varsha Suresh, Timothy Hospedales +1

Sarcasm is a specific type of irony which involves discerning what is said from what is meant. Detecting sarcasm depends not only on the literal content of an utterance but also on…