5 papers
Decompose, Look, and Reason: Reinforced Latent Reasoning for VLMs
Mengdan Zhu, Senhao Cheng, Liang Zhao
Vision-Language Models often struggle with complex visual reasoning due to the visual information loss in textual CoT. Existing methods either add the cost of tool calls or rely on…
Learning User Interests via Reasoning and Distillation for Cross-Domain News Recommendation
Mengdan Zhu, Yufan Zhao, Tao Di +2
News recommendation plays a critical role in online news platforms by helping users discover relevant content. Cross-domain news recommendation further requires inferring user's un…
LatentExplainer: Explaining Latent Representations in Deep Generative Models with Multimodal Large Language Models
Mengdan Zhu, Raasikh Kanjiani, Jiahui Lu +3
Deep generative models like VAEs and diffusion models have advanced various generation tasks by leveraging latent variables to learn data distributions and generate high-quality sa…
Cross-modal RAG: Sub-dimensional Text-to-Image Retrieval-Augmented Generation
Mengdan Zhu, Senhao Cheng, Guangji Bai +2
Text-to-image generation increasingly demands access to domain-specific, fine-grained, and rapidly evolving knowledge that pretrained models cannot fully capture, necessitating the…
Beyond Efficiency: A Systematic Survey of Resource-Efficient Large Language Models
Guangji Bai, Zheng Chai, Chen Ling +11
The burgeoning field of Large Language Models (LLMs), exemplified by sophisticated models like OpenAI's ChatGPT, represents a significant advancement in artificial intelligence. Th…