5 papers
UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation
Zeyuan Yang, Hao-Wei Chen, Xueyang Yu +5
Recent years have seen remarkable progress in unified vision-language models handling both multimodal understanding and generation within a single architecture. While autoregressiv…
GSRM: Generative Speech Reward Model for Speech RLHF
Maohao Shen, Tejas Jayashankar, Osama Hanna +10
Recent advances in speech language models, such as GPT-4o Voice Mode and Gemini Live, have demonstrated promising speech generation capabilities. Nevertheless, the aesthetic natura…
VowelPrompt: Hearing Speech Emotions from Text via Vowel-level Prosodic Augmentation
Yancheng Wang, Osama Hanna, Ruiming Xie +11
Emotion recognition in speech presents a complex multimodal challenge, requiring comprehension of both linguistic content and vocal expressivity, particularly prosodic features suc…
Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens
Zeyuan Yang, Xueyang Yu, Delin Chen +2
Vision-language models (VLMs) excel at multimodal understanding, yet their text-only decoding forces them to verbalize visual reasoning, limiting performance on tasks that demand v…
Get Large Language Models Ready to Speak: A Late-fusion Approach for Speech Generation
Maohao Shen, Shun Zhang, Jilong Wu +5
Large language models (LLMs) have revolutionized natural language processing (NLP) with impressive performance across various text-based tasks. However, the extension of text-domin…