8 papers
SignMAE: Segmentation-Driven Self-Supervised Learning for Sign Language Recognition
Kunyuan Xie, Zhixi Cai, Kalin Stefanov
Subtle hand differences make sign language recognition challenging, yet many existing methods rely on encoders pretrained on generic action datasets that poorly capture such fine-g…
AV-Deepfake1M++: A Large-Scale Audio-Visual Deepfake Benchmark with Real-World Perturbations
Zhixi Cai, Kartik Kuckreja, Shreya Ghosh +5
The rapid surge of text-to-speech and face-voice reenactment models makes video fabrication easier and highly realistic. To encounter this problem, we require datasets that rich in…
Pavlok-Nudge: A Feedback Mechanism for Atomic Behaviour Modification with Snoring Usecase
Md Rakibul Hasan, Shreya Ghosh, Pradyumna Agrawal +3
This paper proposes an atomic behaviour intervention strategy using the Pavlok wearable device. Pavlok utilises beeps, vibration and shocks as a mode of aversion technique to help…
MRAC Track 1: 2nd Workshop on Multimodal, Generative and Responsible Affective Computing
Shreya Ghosh, Zhixi Cai, Abhinav Dhall +3
With the rapid advancements in multimodal generative technology, Affective Computing research has provoked discussion about the potential consequences of AI systems equipped with e…
1M-Deepfakes Detection Challenge
Zhixi Cai, Abhinav Dhall, Shreya Ghosh +4
The detection and localization of deepfake content, particularly when small fake segments are seamlessly mixed with real videos, remains a significant challenge in the field of dig…
Emolysis: A Multimodal Open-Source Group Emotion Analysis and Visualization Toolkit
Shreya Ghosh, Zhixi Cai, Parul Gupta +4
Automatic group emotion recognition plays an important role in understanding complex human-human interaction. This paper introduces, Emolysis, a Python-based, standalone open-sourc…