25 citations · 28 across the 4 of their papers we have counts for
4 papers
Leveraging Unpaired Data for Vision-Language Generative Models via Cycle Consistency
Tianhong Li, Sangnie Bhardwaj, Yonglong Tian +6
Current vision-language generative models rely on expansive corpora of paired image-text data to attain optimal performance and generalization capabilities. However, automatically…
StyleDrop: Text-to-Image Generation in Any Style
Kihyuk Sohn, Nataniel Ruiz, Kimin Lee +11
Pre-trained large text-to-image models synthesize impressive images with an appropriate use of text prompts. However, ambiguities inherent in natural language and out-of-distributi…
SPADE: Self-supervised Pretraining for Acoustic DisEntanglement
John Harvill, Jarred Barber, Arun Nair +1
Self-supervised representation learning approaches have grown in popularity due to the ability to train models on large amounts of unlabeled data and have demonstrated success in d…
Challenges and Opportunities in Multi-device Speech Processing
Gregory Ciccarelli, Jarred Barber, Arun Nair +2
We review current solutions and technical challenges for automatic speech recognition, keyword spotting, device arbitration, speech enhancement, and source localization in multidev…