25 papers
Controllable and Content-Based Recommendations
Fırat Ãncel, Jihoon Jeong, Emiliano Penaloza +3
Traditional recommendation systems rely on latent (dense) representations, making them difficult to interpret and control. We propose the Controllable and Content-Based Recommendat…
PiSAs: Benchmarking Contextual Integrity in Multi-User Agentic Systems
Shubham Gupta, Nazanin Mohammadi Sepahvand, Abhinav Kumar +6
As LLM agents evolve from single-user assistants into shared organizational infrastructure, new privacy risks emerge: inappropriate information may not only be exposed through outp…
Investigating Faithfulness in Large Audio Language Models
Pooneh Mousavi, Lovenya Jain, Mirco Ravanelli +1
Large Audio Language Models (LALMs) integrate audio encoders with pretrained Large Language Models to perform complex multimodal reasoning tasks. While these models can generate Ch…
ALAS: An Automatic Latent Alignment Score for Audio Language Models
Pooneh Mousavi, Yingzhi Wang, Mirco Ravanelli +1
Large Language Models (LLMs) are extended into Speech-LLMs, and the quality of the audio--text alignment they learn affects most downstream Spoken Language Understanding (SLU) beha…
WavSLM: Single-Stream Speech Language Modeling via WavLM Distillation
Luca Della Libera, Cem Subakan, Mirco Ravanelli
Large language models show that simple autoregressive training can yield scalable and coherent generation, but extending this paradigm to speech remains challenging due to the enta…
Exploring Token-Space Manipulation in Latent Audio Tokenizers
Francesco Paissan, Luca Della Libera, Mirco Ravanelli +1
Neural audio codecs provide compact discrete representations for speech generation and manipulation. However, most codecs organize tokens as frame-level sequences, making it diffic…