47 citations · 57 across the 5 of their papers we have counts for
4 papers
Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation
Jiawei Huang, Yi Ren, Rongjie Huang +7
Large diffusion models have been successful in text-to-audio (T2A) synthesis tasks, but they often suffer from common issues such as semantic misalignment and poor temporal consist…
Structure-aware registration network for liver DCE-CT images
Peng Xue, Jingyang Zhang, Lei Ma +7
Image registration of liver dynamic contrast-enhanced computed tomography (DCE-CT) is crucial for diagnosis and image-guided surgical planning of liver cancer. However, intensity v…
Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models
Rongjie Huang, Jiawei Huang, Dongchao Yang +7
Large-scale multimodal generative modeling has created milestones in text-to-image and text-to-video generation. Its application to audio still lags behind for two main reasons: th…
Learning towards Synchronous Network Memorizability and Generalizability for Continual Segmentation across Multiple Sites
Jingyang Zhang, Peng Xue, Ran Gu +7
In clinical practice, a segmentation network is often required to continually learn on a sequential data stream from multiple sites rather than a consolidated set, due to the stora…