7 citations · 21 across the 36 of their papers we have counts for
5 papers · 1 filter
TeMuDance: Contrastive Alignment-Based Textual Control for Music-Driven Dance Generation
Xinran Liu, Diptesh Kanojia, Wenwu Wang +1
Existing music-driven dance generation approaches have achieved strong realism and effective audio-motion alignment. However, they generally lack semantic controllability, making i…
MAGIC-Talk: Motion-aware Audio-Driven Talking Face Generation with Customizable Identity Control
Fatemeh Nazarieh, Zhenhua Feng, Diptesh Kanojia +2
Audio-driven talking face generation has gained significant attention for applications in digital media and virtual avatars. While recent methods improve audio-lip synchronization,…
TRACE: Textual Relevance Augmentation and Contextual Encoding for Multimodal Hate Detection
Girish A. Koushik, Helen Treharne, Aditya Joshi +1
Social media memes are a challenging domain for hate detection because they intertwine visual and textual cues into culturally nuanced messages. To tackle these challenges, we intr…
Towards a Robust Framework for Multimodal Hate Detection: A Study on Video vs. Image-based Content
Girish A. Koushik, Diptesh Kanojia, Helen Treharne
Social media platforms enable the propagation of hateful content across different modalities such as textual, auditory, and visual, necessitating effective detection methods. While…
PortraitTalk: Towards Customizable One-Shot Audio-to-Talking Face Generation
Fatemeh Nazarieh, Zhenhua Feng, Diptesh Kanojia +2
Audio-driven talking face generation is a challenging task in digital communication. Despite significant progress in the area, most existing methods concentrate on audio-lip synchr…