4 papers · 1 filter
Acoustic Cue Alignment in Audio Language Models for Speech Emotion Recognition
Iosif Tsangko, Andreas Triantafyllopoulos, Björn W. Schuller
Instruction-following audio language models (ALMs) can be augmented with explicit acoustic cues, yet it remains unclear whether such cues are used in a grounded way when the raw au…
Computer Audition: From Task-Specific Machine Learning to Foundation Models
Andreas Triantafyllopoulos, Iosif Tsangko, Alexander Gebhard +3
Foundation models (FMs) are increasingly spearheading recent advances on a variety of tasks that fall under the purview of computer audition -- the use of machines to understand so…
DFingerNet: Noise-Adaptive Speech Enhancement for Hearing Aids
Iosif Tsangko, Andreas Triantafyllopoulos, Michael Müller +2
The DeepFilterNet (DFN) architecture was recently proposed as a deep learning model suited for hearing aid devices. Despite its competitive performance on numerous benchmarks, it s…
Abusive Speech Detection in Indic Languages Using Acoustic Features
Anika A. Spiesberger, Andreas Triantafyllopoulos, Iosif Tsangko +1
Abusive content in online social networks is a well-known problem that can cause serious psychological harm and incite hatred. The ability to upload audio data increases the import…