2 papers
cs.LG2026
From Perception to Punchline: Empowering VLM with the Art of In-the-wild Meme
Xueyan Li, Yingyi Xue, Mengjie Jiang +2
Generating humorous memes is a challenging multimodal task that moves beyond direct image-to-caption supervision. It requires a nuanced reasoning over visual content, contextual cu…
cs.SD2025
Generative Multi-modal Feedback for Singing Voice Synthesis Evaluation
Xueyan Li, Yuxin Wang, Mengjie Jiang +4
Singing voice synthesis (SVS) has advanced significantly, enabling models to generate vocals with accurate pitch and consistent style. As these capabilities improve, the need for r…