4 papers · 1 filter
To Be Multimodal or Not to Be: Query-Adaptive Audio-Visual Person Retrieval via Active Modality Detection
Erfan Loweimi, Mengjie Qian, Kate Knill +7
When retrieving a person from a video archive by voice and face, should the system be multimodal or not? In real-world broadcast archives, unlike curated benchmarks, a target may b…
Predicting Psychological Well-Being from Spontaneous Speech using LLMs
Erfan Loweimi, Sofia de la Fuente Garcia, Saturnino Luz
We investigate the use of Large Language Models (LLMs) for zero-shot prediction of Ryff Psychological Well-Being (PWB) scores from spontaneous speech. Using a few minutes of voice…
Can We Trust LLMs for Mental Health Screening? Consistency, ASR Robustness, and Evidence Faithfulness
Erfan Loweimi, Sofia de la Fuente Garcia, Samira Loveymi +2
LLMs can estimate Hospital Anxiety and Depression Scale (HADS) scores from speech in a zero-shot manner, but clinical deployment requires reliability across three dimensions: intra…
Zero-shot Audio Topic Reranking using Large Language Models
Mengjie Qian, Rao Ma, Adian Liusie +3
Multimodal Video Search by Examples (MVSE) investigates using video clips as the query term for information retrieval, rather than the more traditional text query. This enables far…