1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.CV2024
EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation
Shuhao Han, Haotian Fan, Jiachen Fu +8
Recently, Text-to-Image (T2I) generation models have achieved significant advancements. Correspondingly, many automated metrics have emerged to evaluate the image-text alignment ca…
cs.CV2024
Mixture of Rationale: Multi-Modal Reasoning Mixture for Visual Question Answering
Tao Li, Linjun Shou, Xuejun Liu
Zero-shot visual question answering (VQA) is a challenging task that requires reasoning across modalities. While some existing methods rely on a single rationale within the Chain o…
cs.SD2023★ 1 cited
Interpretable Style Transfer for Text-to-Speech with ControlVAE and Diffusion Bridge
Wenhao Guan, Tao Li, Yishuang Li +3
With the demand for autonomous control and personalized speech generation, the style control and transfer in Text-to-Speech (TTS) is becoming more and more important. In this paper…