2 citations · 4 across the 8 of their papers we have counts for
7 papers · 1 filter
Semantic Motion Anchors: Bridging Motion and Meaning in Co-Speech Gestures
Varsha Suresh, Mohammad Mahdi Abootorabi, Mohamed Salman +5
Learning a shared representation between spoken text and gesture is central to co-speech gesture retrieval, synthesis, and understanding, but remains challenging for semantically m…
Generation-Step-Aware Framework for Cross-Modal Representation and Control in Multilingual Speech-Text Models
Toshiki Nakai, Varsha Suresh, Vera Demberg
Multilingual speech-text models rely on cross-modal language alignment to transfer knowledge between speech and text, but it remains unclear whether this reflects shared computatio…
System-Mediated Attention Imbalances Make Vision-Language Models Say Yes
Tsan Tsai Chan, Varsha Suresh, Anisha Saha +2
Vision-language model (VLM) hallucination is commonly linked to imbalanced allocation of attention across input modalities: system, image and text. However, existing mitigation str…
Modeling Turn-Taking with Semantically Informed Gestures
Varsha Suresh, M. Hamza Mughal, Christian Theobalt +1
In conversation, humans use multimodal cues, such as speech, gestures, and gaze, to manage turn-taking. While linguistic and acoustic features are informative, gestures provide com…
Synthetic Data Augmentation for Cross-domain Implicit Discourse Relation Recognition
Frances Yung, Varsha Suresh, Zaynab Reza +2
Implicit discourse relation recognition (IDRR) -- the task of identifying the implicit coherence relation between two text spans -- requires deep semantic understanding. Recent stu…
Enhancing Spoken Discourse Modeling in Language Models Using Gestural Cues
Varsha Suresh, M. Hamza Mughal, Christian Theobalt +1
Research in linguistics shows that non-verbal cues, such as gestures, play a crucial role in spoken discourse. For example, speakers perform hand gestures to indicate topic shifts,…