4 papers · 1 filter
TokenVerse: Versatile Multi-concept Personalization in Token Modulation Space
Daniel Garibi, Shahar Yadin, Roni Paiss +6
We present TokenVerse -- a method for multi-concept personalization, leveraging a pre-trained text-to-image diffusion model. Our framework can disentangle complex visual elements a…
SpeedNet: Learning the Speediness in Videos
Sagie Benaim, Ariel Ephrat, Oran Lang +5
We wish to automatically predict the "speediness" of moving objects in videos---whether they move faster, at, or slower than their "natural" speed. The core component in our approa…
Dynamic Temporal Alignment of Speech to Lips
Tavi Halperin, Ariel Ephrat, Shmuel Peleg
Many speech segments in movies are re-recorded in a studio during postproduction, to compensate for poor sound quality as recorded on location. Manual alignment of the newly-record…
Improved Speech Reconstruction from Silent Video
Ariel Ephrat, Tavi Halperin, Shmuel Peleg
Speechreading is the task of inferring phonetic information from visually observed articulatory facial movements, and is a notoriously difficult task for humans to perform. In this…