2 papers
cs.SD2026
Investigating the Viability of Employing Multi-modal Large Language Models in the Context of Audio Deepfake Detection
Akanksha Chuchra, Shukesh Reddy, Sudeepta Mishra +2
While Vision-Language Models (VLMs) and Multimodal Large Language Models (MLLMs) have shown strong generalisation in detecting image and video deepfakes, their use for audio deepfa…
cs.CV2025
SynchroRaMa : Lip-Synchronized and Emotion-Aware Talking Face Generation via Multi-Modal Emotion Embedding
Phyo Thet Yee, Dimitrios Kollias, Sudeepta Mishra +1
Audio-driven talking face generation has received growing interest, particularly for applications requiring expressive and natural human-avatar interaction. However, most existing…