7 papers
Acoustic Cue Alignment in Audio Language Models for Speech Emotion Recognition
Iosif Tsangko, Andreas Triantafyllopoulos, Björn W. Schuller
Instruction-following audio language models (ALMs) can be augmented with explicit acoustic cues, yet it remains unclear whether such cues are used in a grounded way when the raw au…
A Pilot Study on Curator-Guided Multilingual Art Description for Blind and Low-Vision Audiences with Small Vision-Language Models
Iosif Tsangko, Andreas Triantafyllopoulos, George Margetis +2
Blind and low-vision (BLV) audiences remain underserved by visual art descriptions, particularly across languages and in museum settings where privacy and intellectual-property con…
Computer Audition: From Task-Specific Machine Learning to Foundation Models
Andreas Triantafyllopoulos, Iosif Tsangko, Alexander Gebhard +3
Foundation models (FMs) are increasingly spearheading recent advances on a variety of tasks that fall under the purview of computer audition -- the use of machines to understand so…
Reading Smiles: Proxy Bias in Foundation Models for Facial Emotion Recognition
Iosif Tsangko, Andreas Triantafyllopoulos, Adem Abdelmoula +2
Foundation Models (FMs) are rapidly transforming Affective Computing (AC), with Vision Language Models (VLMs) now capable of recognising emotions in zero shot settings. This paper…
MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge
Xin Jing, Jiadong Wang, Iosif Tsangko +2
Although speech emotion recognition (SER) has advanced significantly with deep learning, annotation remains a major hurdle. Human annotation is not only costly but also subject to…
DFingerNet: Noise-Adaptive Speech Enhancement for Hearing Aids
Iosif Tsangko, Andreas Triantafyllopoulos, Michael Müller +2
The DeepFilterNet (DFN) architecture was recently proposed as a deep learning model suited for hearing aid devices. Despite its competitive performance on numerous benchmarks, it s…