14 papers
CoughPhase-CLR: Designing an acoustics-informed foundation model for coughing sound classification
Marius Moldovan, Anton Batliner, Thomas M. Berghaus +2
In this work, we introduce CoughPhase-CLR, a self-supervised learning framework designed to leverage the physiological phases of a cough for robust representation learning. Unlike…
Acoustic Cue Alignment in Audio Language Models for Speech Emotion Recognition
Iosif Tsangko, Andreas Triantafyllopoulos, Björn W. Schuller
Instruction-following audio language models (ALMs) can be augmented with explicit acoustic cues, yet it remains unclear whether such cues are used in a grounded way when the raw au…
A Pilot Study on Curator-Guided Multilingual Art Description for Blind and Low-Vision Audiences with Small Vision-Language Models
Iosif Tsangko, Andreas Triantafyllopoulos, George Margetis +2
Blind and low-vision (BLV) audiences remain underserved by visual art descriptions, particularly across languages and in museum settings where privacy and intellectual-property con…
CoarseSoundNet: Building a reliable model for ecological soundscape analysis
Alexander Gebhard, Andreas Triantafyllopoulos, Dominik Arend +4
A soundscape is composed of three types of sound: biophony (sounds made by animals), geophony (natural abiotic sounds) and anthropophony (sounds made by humans). A key research que…
A conceptual framework for learning to listen by reward: Curiosity-driven search for novel sources
Andreas Triantafyllopoulos, Jakub Šťastný, Alexios Terpinas +3
Reinforcement learning is a powerful learning paradigm that has spearheaded progress in numerous domains. Its core promise lies in learning through high-level goals without the nee…
How Class Ontology and Data Scale Affect Audio Transfer Learning
Manuel Milling, Andreas Triantafyllopoulos, Alexander Gebhard +2
Transfer learning is a crucial concept within deep learning that allows artificial neural networks to benefit from a large pre-training data basis when confronted with a task of li…