9 papers
Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing
Iftach Shoham, Tali Dror, Oren Gal +3
Speech recordings often contain missing, corrupted, or incorrect regions that must be reconstructed or modified without re-synthesizing the entire utterance. Speech inpainting rest…
Where Vision Becomes Text: Locating the OCR Routing Bottleneck in Vision-Language Models
Jonathan Steinberg, Oren Gal
Vision-language models (VLMs) can read text from images, but where does this optical character recognition (OCR) information enter the language processing stream? We investigate th…
Split and Conquer Partial Deepfake Speech
Inbal Rimon, Oren Gal, Haim Permuter
Partial deepfake speech detection requires identifying manipulated regions that may occur within short temporal portions of an otherwise bona fide utterance, making the task partic…
Token-Based Audio Inpainting via Discrete Diffusion
Tali Dror, Iftach Shoham, Moshe Buchris +4
Audio inpainting seeks to restore missing segments in degraded recordings. Previous diffusion-based methods exhibit impaired performance when the missing region is large. We introd…
Unmasking Deepfakes: Leveraging Augmentations and Features Variability for Deepfake Speech Detection
Inbal Rimon, Oren Gal, Haim Permuter
Deepfake speech detection presents a growing challenge as generative audio technologies continue to advance. We propose a hybrid training framework that advances detection performa…
Autonomous Oil Spill Response Through Liquid Neural Trajectory Modeling and Coordinated Marine Robotics
Hadas C. Kuzmenko, David Ehevich, Oren Gal
Marine oil spills pose grave environmental and economic risks, threatening marine ecosystems, coastlines, and dependent industries. Predicting and managing oil spill trajectories i…