collaborators
Showing cs.SDShow all

5 papers · 1 filter

cs.SD2025

AudioToolAgent: An Agentic Framework for Audio-Language Models

Gijs Wijngaard, Elia Formisano, Michel Dumontier +1

Large Audio-Language Models (LALMs) perform well on audio understanding tasks but lack multistep reasoning and tool-calling found in recent Large Language Models (LLMs). This paper…

cs.SD2025

Data-Balanced Curriculum Learning for Audio Question Answering

Gijs Wijngaard, Elia Formisano, Michele Esposito +1

Audio question answering (AQA) requires models to understand acoustic content and perform complex reasoning. Current models struggle with dataset imbalances and unstable training d…

cs.SD2025

AudSemThinker: Enhancing Audio-Language Models through Reasoning over Semantics of Sound

Gijs Wijngaard, Elia Formisano, Michele Esposito +1

Audio-language models have shown promising results in various sound understanding tasks, yet they remain limited in their ability to reason over the fine-grained semantics of sound…

cs.SD2024

Audio-Language Datasets of Scenes and Events: A Survey

Gijs Wijngaard, Elia Formisano, Michele Esposito +1

Audio-language models (ALMs) generate linguistic descriptions of sound-producing events and scenes. Advances in dataset creation and computational power have led to significant pro…

cs.SD2024

ACES: Evaluating Automated Audio Captioning Models on the Semantics of Sounds

Gijs Wijngaard, Elia Formisano, Bruno L. Giordano +1

Automated Audio Captioning is a multimodal task that aims to convert audio content into natural language. The assessment of audio captioning systems is typically based on quantitat…