From the 1 of 8 linked papers with an AI index.
8 papers
CARE: A Multimodal Corpus for Studying Speech and Non-Verbal Communication Across Multiple Medical Conditions
David Gimeno-Gómez, Catarina Botelho, Carlos-D. MartÃnez-Hinarejos +2
The paper introduces CARE v1.0, a curated multimodal English dataset of about 144 hours of short video interviews from 612 participants covering 12 medical conditions and a control…
Cross-Modal Masking for Robust Silent Speech Synthesis Using sEMG and Lipreading
Eder del Blanco, David Gimeno-Gómez, Eva Navas +2
Speech restoration through silent speech interfaces (SSIs) has emerged as a promising assistive technology for individuals with impaired or absent laryngeal voice production. Among…
MEMEWEAVER: Inter-Meme Graph Reasoning for Sexism and Misogyny Detection
Paolo Italiani, David Gimeno-Gomez, Luca Ragazzi +2
Women are twice as likely as men to face online harassment due to their gender. Despite recent advances in multimodal content moderation, most approaches still overlook the social…
On the Relevance of Clinical Assessment Tasks for the Automatic Detection of Parkinson's Disease Medication State from Speech
David Gimeno-Gómez, Rubén Solera-Ureña, Anna Pompili +5
The automatic identification of medication states of Parkinson's disease (PD) patients can assist clinicians in monitoring and scheduling personalized treatments, as well as studyi…
Tailored Design of Audio-Visual Speech Recognition Models using Branchformers
David Gimeno-Gómez, Carlos-D. MartÃnez-Hinarejos
Recent advances in Audio-Visual Speech Recognition (AVSR) have led to unprecedented achievements in the field, improving the robustness of this type of system in adverse, noisy env…
Evaluation of End-to-End Continuous Spanish Lipreading in Different Data Conditions
David Gimeno-Gómez, Carlos-D. MartÃnez-Hinarejos
Visual speech recognition remains an open research problem where different challenges must be considered by dispensing with the auditory sense, such as visual ambiguities, the inte…