1 paper · 1 filter
Zhiyong Ma, Jiahao Chen, Qingyuan Chuai +1
Multi-modal generation struggles to ensure thematic coherence and style consistency. Semantically, existing methods suffer from cross-modal mismatch and lack explicit modeling of c…