3 papers
cs.LG2026
From Perception to Punchline: Empowering VLM with the Art of In-the-wild Meme
Xueyan Li, Yingyi Xue, Mengjie Jiang +2
Generating humorous memes is a challenging multimodal task that moves beyond direct image-to-caption supervision. It requires a nuanced reasoning over visual content, contextual cu…
cs.SD2025
Generative Multi-modal Feedback for Singing Voice Synthesis Evaluation
Xueyan Li, Yuxin Wang, Mengjie Jiang +4
Singing voice synthesis (SVS) has advanced significantly, enabling models to generate vocals with accurate pitch and consistent style. As these capabilities improve, the need for r…
cs.LG2025
Empowering LLMs in Decision Games through Algorithmic Data Synthesis
Haolin Wang, Xueyan Li, Yazhe Niu +2
Large Language Models (LLMs) have exhibited impressive capabilities across numerous domains, yet they often struggle with complex reasoning and decision-making tasks. Decision-maki…