3 papers
cs.LG2026
TokenMapper: A Step Toward Interoperable Speech Token Translation
Tal Kozakov, Tal Rosenwein, Eliya Nachmani
Neural audio codecs discretize speech into token sequences, but the resulting token spaces differ in vocabulary and codebook structure, preventing direct communication across model…
eess.AS2024
RevRIR: Joint Reverberant Speech and Room Impulse Response Embedding using Contrastive Learning with Application to Room Shape Classification
Jacob Bitterman, Daniel Levi, Hilel Hagai Diamandi +2
This paper focuses on room fingerprinting, a task involving the analysis of an audio recording to determine the specific volume and shape of the room in which it was captured. Whil…
cs.CV2016
Learning a Metric Embedding for Face Recognition using the Multibatch Method
Oren Tadmor, Yonatan Wexler, Tal Rosenwein +2
This work is motivated by the engineering task of achieving a near state-of-the-art face recognition on a minimal computing budget running on an embedded system. Our main technical…