collaborators

8 papers

cs.SD2026

Are Audio-Language Models Listening? Audio-Specialist Heads for Adaptive Audio Steering

Neta Glazer, Lenny Aharon, Ethan Fetaya

Multimodal large language models can exhibit text dominance, over-relying on linguistic priors instead of grounding predictions in non-text inputs. One example is large audio-langu…

cs.CV2025

Questioning the Stability of Visual Question Answering

Amir Rosenfeld, Neta Glazer, Ethan Fetaya

Visual Language Models (VLMs) have achieved remarkable progress, yet their reliability under small, meaning-preserving input changes remains poorly understood. We present the first…

cs.LG2025

Multi Task Inverse Reinforcement Learning for Common Sense Reward

Neta Glazer, Aviv Navon, Aviv Shamsian +1

One of the challenges in applying reinforcement learning in a complex real-world environment lies in providing the agent with a sufficiently detailed reward function. Any misalignm…

eess.AS2025

Drax: Speech Recognition with Discrete Flow Matching

Aviv Navon, Aviv Shamsian, Neta Glazer +4

Diffusion and flow-based non-autoregressive (NAR) models have shown strong promise in large language modeling, however, their potential for automatic speech recognition (ASR) remai…

cs.SD2025

Beyond Transcription: Mechanistic Interpretability in ASR

Neta Glazer, Yael Segal-Feldman, Hilit Segev +6

Interpretability methods have recently gained significant attention, particularly in the context of large language models, enabling insights into linguistic representations, error…

cs.SD2025

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching

Neta Glazer, Aviv Navon, Yael Segal +6

Recent advances in Text-to-Speech (TTS) have enabled highly natural speech synthesis, yet integrating speech with complex background environments remains challenging. We introduce…