activity
20242026
collaborators

5 papers

cs.CL2026

Decompose, Look, and Reason: Reinforced Latent Reasoning for VLMs

Mengdan Zhu, Senhao Cheng, Liang Zhao

Vision-Language Models often struggle with complex visual reasoning due to the visual information loss in textual CoT. Existing methods either add the cost of tool calls or rely on…

cs.CL2026

Learning User Interests via Reasoning and Distillation for Cross-Domain News Recommendation

Mengdan Zhu, Yufan Zhao, Tao Di +2

News recommendation plays a critical role in online news platforms by helping users discover relevant content. Cross-domain news recommendation further requires inferring user's un…

cs.LG2025

LatentExplainer: Explaining Latent Representations in Deep Generative Models with Multimodal Large Language Models

Mengdan Zhu, Raasikh Kanjiani, Jiahui Lu +3

Deep generative models like VAEs and diffusion models have advanced various generation tasks by leveraging latent variables to learn data distributions and generate high-quality sa…

cs.CV2025

Cross-modal RAG: Sub-dimensional Text-to-Image Retrieval-Augmented Generation

Mengdan Zhu, Senhao Cheng, Guangji Bai +2

Text-to-image generation increasingly demands access to domain-specific, fine-grained, and rapidly evolving knowledge that pretrained models cannot fully capture, necessitating the…

cs.LG2024

Beyond Efficiency: A Systematic Survey of Resource-Efficient Large Language Models

Guangji Bai, Zheng Chai, Chen Ling +11

The burgeoning field of Large Language Models (LLMs), exemplified by sophisticated models like OpenAI's ChatGPT, represents a significant advancement in artificial intelligence. Th…