47 citations · 53 across the 4 of their papers we have counts for
4 papers
Extending Multi-modal Contrastive Representations
Zehan Wang, Ziang Zhang, Luping Liu +4
Multi-modal contrastive representation (MCR) of more than three modalities is critical in multi-modal learning. Although recent methods showcase impressive achievements, the high d…
Detector Guidance for Multi-Object Text-to-Image Generation
Luping Liu, Zijian Zhang, Yi Ren +3
Diffusion models have demonstrated impressive performance in text-to-image generation. They utilize a text encoder and cross-attention blocks to infuse textual information into ima…
Make-A-Voice: Unified Voice Synthesis With Discrete Representation
Rongjie Huang, Chunlei Zhang, Yongqi Wang +7
Various applications of voice synthesis have been developed independently despite the fact that they generate "voice" as output in common. In addition, the majority of voice synthe…
Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models
Rongjie Huang, Jiawei Huang, Dongchao Yang +7
Large-scale multimodal generative modeling has created milestones in text-to-image and text-to-video generation. Its application to audio still lags behind for two main reasons: th…