collaborators

5 papers

cs.MM2025

Quality Over Quantity? LLM-Based Curation for a Data-Efficient Audio-Video Foundation Model

Ali Vosoughi, Dimitra Emmanouilidou, Hannes Gamper

Integrating audio and visual data for training multimodal foundational models remains a challenge. The Audio-Video Vector Alignment (AVVA) framework addresses this by considering A…

cs.SD2025

FOA Tokenizer: Low-bitrate Neural Codec for First Order Ambisonics with Spatial Consistency Loss

Parthasaarathy Sudarsanam, Sebastian Braun, Hannes Gamper

Neural audio codecs have been widely studied for mono and stereo signals, but spatial audio remains largely unexplored. We present the first discrete neural spatial audio codec for…

eess.AS2025

SALAD-VAE: Semantic Audio Compression with Language-Audio Distillation

Sebastian Braun, Hannes Gamper, Dimitra Emmanouilidou

Modern generative and multimodal models increasingly rely on compact latent representations that trade and balance semantic richness with high-fidelity reconstruction. We introduce…

eess.AS2025

Make Some Noise: Towards LLM audio reasoning and generation using sound tokens

Shivam Mehta, Nebojsa Jojic, Hannes Gamper

Integrating audio comprehension and generation into large language models (LLMs) remains challenging due to the continuous nature of audio and the resulting high sampling rates. He…

eess.AS2025

Distillation and Pruning for Scalable Self-Supervised Representation-Based Speech Quality Assessment

Benjamin Stahl, Hannes Gamper

In this paper, we investigate distillation and pruning methods to reduce model size for non-intrusive speech quality assessment based on self-supervised representations. Our experi…