4 citations · 6 across the 6 of their papers we have counts for
6 papers · 1 filter
To Be Multimodal or Not to Be: Query-Adaptive Audio-Visual Person Retrieval via Active Modality Detection
Erfan Loweimi, Mengjie Qian, Kate Knill +7
When retrieving a person from a video archive by voice and face, should the system be multimodal or not? In real-world broadcast archives, unlike curated benchmarks, a target may b…
Predicting Psychological Well-Being from Spontaneous Speech using LLMs
Erfan Loweimi, Sofia de la Fuente Garcia, Saturnino Luz
We investigate the use of Large Language Models (LLMs) for zero-shot prediction of Ryff Psychological Well-Being (PWB) scores from spontaneous speech. Using a few minutes of voice…
Can We Trust LLMs for Mental Health Screening? Consistency, ASR Robustness, and Evidence Faithfulness
Erfan Loweimi, Sofia de la Fuente Garcia, Samira Loveymi +2
LLMs can estimate Hospital Anxiety and Depression Scale (HADS) scores from speech in a zero-shot manner, but clinical deployment requires reliability across three dimensions: intra…
Zero-shot Audio Topic Reranking using Large Language Models
Mengjie Qian, Rao Ma, Adian Liusie +3
Multimodal Video Search by Examples (MVSE) investigates using video clips as the query term for information retrieval, rather than the more traditional text query. This enables far…
On the Usefulness of Self-Attention for Automatic Speech Recognition with Transformers
Shucong Zhang, Erfan Loweimi, Peter Bell +1
Self-attention models such as Transformers, which can capture temporal relationships without being limited by the distance between events, have given competitive speech recognition…
Stochastic Attention Head Removal: A simple and effective method for improving Transformer Based ASR Models
Shucong Zhang, Erfan Loweimi, Peter Bell +1
Recently, Transformer based models have shown competitive automatic speech recognition (ASR) performance. One key factor in the success of these models is the multi-head attention…