47 citations · 102 across the 4 of their papers we have counts for
1 paper · 1 filter
Rongjie Huang, Jiawei Huang, Dongchao Yang +7
Large-scale multimodal generative modeling has created milestones in text-to-image and text-to-video generation. Its application to audio still lags behind for two main reasons: th…