2 papers
cs.CV2026
Intermediate Text Representation Guided Text-to-Image Generation for Enhancing One-and-Only Alignment
Soyoun Won, Aryan Yazdan Parast, Basim Azam +2
Text-to-image (T2I) diffusion models often fail to faithfully render explicit textual descriptions, instead defaulting to strongly learned visual priors due to a phenomenon referre…
cs.SD2026
Disentangling Reasoning in Large Audio-Language Models for Ambiguous Emotion Prediction
Xiaofeng Yu, Jiaheng Dong, Jean Honorio +3
Speech emotion recognition plays an important role in various applications. However, most existing approaches predict a single emotion label, oversimplifying the inherently ambiguo…