4 papers
Beyond Transcription: Mechanistic Interpretability in ASR
Neta Glazer, Yael Segal-Feldman, Hilit Segev +6
Interpretability methods have recently gained significant attention, particularly in the context of large language models, enabling insights into linguistic representations, error…
UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching
Neta Glazer, Aviv Navon, Yael Segal +6
Recent advances in Text-to-Speech (TTS) have enabled highly natural speech synthesis, yet integrating speech with complex background environments remains challenging. We introduce…
Make It Count: Text-to-Image Generation with an Accurate Number of Objects
Lital Binyamin, Yoad Tewel, Hilit Segev +3
Despite the unprecedented success of text-to-image diffusion models, controlling the number of depicted objects using text is surprisingly hard. This is important for various appli…
Lay-A-Scene: Personalized 3D Object Arrangement Using Text-to-Image Priors
Ohad Rahamim, Hilit Segev, Idan Achituve +3
Generating 3D visual scenes is at the forefront of visual generative AI, but current 3D generation techniques struggle with generating scenes with multiple high-resolution objects.…