most citedMoFusion: A Framework for Denoising-Diffusion-based Motion Synthesis

5 citations · 5 across the 5 of their papers we have counts for

collaborators

8 papers

cs.CV20265 cited

MoFusion: A Framework for Denoising-Diffusion-based Motion Synthesis

Rishabh Dabral, Muhammad Hamza Mughal, Vladislav Golyanik +1

Conventional methods for human motion synthesis are either deterministic or struggle with the trade-off between motion diversity and motion quality. In response to these limitation…

cs.CV2026

ReFree: Towards Realistic Co-Speech Video Generation via Reward-Free RL and Multilevel Speech Guidance

Salaheldin Mohamed, M. Hamza Mughal, Rishabh Dabral +1

Speech-driven talking character animation seeks to generate life-like portrait videos that convey natural conversation behavior, aligning facial motion with spoken audio. Although…

cs.CL2026

Semantic Motion Anchors: Bridging Motion and Meaning in Co-Speech Gestures

Varsha Suresh, Mohammad Mahdi Abootorabi, Mohamed Salman +5

Learning a shared representation between spoken text and gesture is central to co-speech gesture retrieval, synthesis, and understanding, but remains challenging for semantically m…

cs.CV2026

Towards Reliable Human Evaluations in Gesture Generation: Insights from a Community-Driven State-of-the-Art Benchmark

Rajmund Nagy, Hendric Voss, Thanh Hoang-Minh +18

We review human evaluation practices in automatic, speech-driven 3D gesture generation and find a lack of standardisation and frequent use of flawed experimental setups. This leads…

cs.CV2026

MIBURI: Towards Expressive Interactive Gesture Synthesis

M. Hamza Mughal, Rishabh Dabral, Vera Demberg +1

Embodied Conversational Agents (ECAs) aim to emulate human face-to-face interaction through speech, gestures, and facial expressions. Current large language model (LLM)-based conve…

cs.CL2026

Modeling Turn-Taking with Semantically Informed Gestures

Varsha Suresh, M. Hamza Mughal, Christian Theobalt +1

In conversation, humans use multimodal cues, such as speech, gestures, and gaze, to manage turn-taking. While linguistic and acoustic features are informative, gestures provide com…