4 papers · 1 filter
Are Audio-Language Models Listening? Audio-Specialist Heads for Adaptive Audio Steering
Neta Glazer, Lenny Aharon, Ethan Fetaya
Multimodal large language models can exhibit text dominance, over-relying on linguistic priors instead of grounding predictions in non-text inputs. One example is large audio-langu…
Beyond Transcription: Mechanistic Interpretability in ASR
Neta Glazer, Yael Segal-Feldman, Hilit Segev +6
Interpretability methods have recently gained significant attention, particularly in the context of large language models, enabling insights into linguistic representations, error…
UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching
Neta Glazer, Aviv Navon, Yael Segal +6
Recent advances in Text-to-Speech (TTS) have enabled highly natural speech synthesis, yet integrating speech with complex background environments remains challenging. We introduce…
Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes
Neta Glazer, David Chernin, Idan Achituve +2
Recent advancements in Text-to-Speech (TTS) models, particularly in voice cloning, have intensified the demand for adaptable and efficient deepfake detection methods. As TTS system…