activity
20242026
collaborators

5 papers

cs.CV2026

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation

Zeyuan Yang, Hao-Wei Chen, Xueyang Yu +5

Recent years have seen remarkable progress in unified vision-language models handling both multimodal understanding and generation within a single architecture. While autoregressiv…

cs.SD2026

GSRM: Generative Speech Reward Model for Speech RLHF

Maohao Shen, Tejas Jayashankar, Osama Hanna +10

Recent advances in speech language models, such as GPT-4o Voice Mode and Gemini Live, have demonstrated promising speech generation capabilities. Nevertheless, the aesthetic natura…

cs.CL2026

VowelPrompt: Hearing Speech Emotions from Text via Vowel-level Prosodic Augmentation

Yancheng Wang, Osama Hanna, Ruiming Xie +11

Emotion recognition in speech presents a complex multimodal challenge, requiring comprehension of both linguistic content and vocal expressivity, particularly prosodic features suc…

cs.CV2025

Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Zeyuan Yang, Xueyang Yu, Delin Chen +2

Vision-language models (VLMs) excel at multimodal understanding, yet their text-only decoding forces them to verbalize visual reasoning, limiting performance on tasks that demand v…

cs.CL2024

Get Large Language Models Ready to Speak: A Late-fusion Approach for Speech Generation

Maohao Shen, Shun Zhang, Jilong Wu +5

Large language models (LLMs) have revolutionized natural language processing (NLP) with impressive performance across various text-based tasks. However, the extension of text-domin…