4 papers
VISTAR:A User-Centric and Role-Driven Benchmark for Text-to-Image Evaluation
Kaiyuan Jiang, Ruoxi Sun, Ying Cao +4
We present VISTAR, a user-centric, multi-dimensional benchmark for text-to-image (T2I) evaluation that addresses the limitations of existing metrics. VISTAR introduces a two-tier h…
EDTC: enhance depth of text comprehension in automated audio captioning
Liwen Tan, Yin Cao, Yi Zhou
Modality discrepancies have perpetually posed significant challenges within the realm of Automated Audio Captioning (AAC) and across all multi-modal domains. Facilitating models in…
HieraFashDiff: Hierarchical Fashion Design with Multi-stage Diffusion Models
Zhifeng Xie, Hao Li, Huiming Ding +3
Fashion design is a challenging and complex process.Recent works on fashion generation and editing are all agnostic of the actual fashion design process, which limits their usage i…
Balanced SNR-Aware Distillation for Guided Text-to-Audio Generation
Bingzhi Liu, Yin Cao, Haohe Liu +1
Diffusion models have demonstrated promising results in text-to-audio generation tasks. However, their practical usability is hindered by slow sampling speeds, limiting their appli…