21 citations · 23 across the 3 of their papers we have counts for
3 papers
cs.SD2024★ 2 cited
Text-to-Audio Generation Synchronized with Videos
Shentong Mo, Jing Shi, Yapeng Tian
In recent times, the focus on text-to-audio (TTA) generation has intensified, as researchers strive to synthesize audio from textual descriptions. However, most existing methods, t…
cs.CV2023★ 21 cited
InstantBooth: Personalized Text-to-Image Generation without Test-Time Finetuning
Jing Shi, Wei Xiong, Zhe Lin +1
Recent advances in personalized image generation allow a pre-trained text-to-image model to learn a new concept from a set of images. However, existing personalization approaches u…
cs.CL2023
Matching-based Term Semantics Pre-training for Spoken Patient Query Understanding
Zefa Hu, Xiuyi Chen, Haoran Wu +5
Medical Slot Filling (MSF) task aims to convert medical queries into structured information, playing an essential role in diagnosis dialogue systems. However, the lack of sufficien…