9 papers
OLIVE: View-Augmented Latent Prediction with Waveform Reconstruction for Speech SSL
Karl El Hajal, Mathew Magimai. -Doss
We propose Online Latent prediction with Invariant Views and rEconstruction (OLIVE), a self-supervised speech representation learning framework that jointly optimizes analysis and…
Assessment of Personality Dimensions Across Situations in Dyadic Role-Play Scenarios
Alice Zhang, Skanda Muralidhar, Daniel Gatica-Perez +1
Prior research indicates that users prefer assistive technologies whose personalities align with their own. This has sparked interest in automatic personality perception (APP), whi…
Toward using Speech to Sense Student Emotion in Remote Learning Environments
Sargam Vyas, Bogdan Vlasenko, André Mayoraz +3
With advancements in multimodal communication technologies, remote learning environments such as, distance universities are increasing. Remote learning typically happens asynchrono…
Multilingual Hidden Prompt Injection Attacks on LLM-Based Academic Reviewing
Panagiotis Theocharopoulos, Ajinkya Kulkarni, Mathew Magimai. -Doss
Large language models (LLMs) are increasingly considered for use in high-impact workflows, including academic peer review. However, LLMs are vulnerable to document-level hidden pro…
Unveiling Audio Deepfake Origins: A Deep Metric learning And Conformer Network Approach With Ensemble Fusion
Ajinkya Kulkarni, Sandipana Dowerah, Tanel Alumae +1
Audio deepfakes are acquiring an unprecedented level of realism with advanced AI. While current research focuses on discerning real speech from spoofed speech, tracing the source s…
Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech
Karl El Hajal, Enno Hermann, Sevada Hovsepyan +1
Automatic speech recognition (ASR) systems struggle with dysarthric speech due to high inter-speaker variability and slow speaking rates. To address this, we explore dysarthric-to-…