activity
20242026
most citedGLOV: Guided Large Language Models as Implicit Optimizers for Vision Language Models

1 citations · 2 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CL2026

MemDefrag: Latent Memory Defragmentation for Large Language Models

Ruiyi Yan, Zhuoyuan Mao, Yiwen Guo

Latent memory, which stores past knowledge fragments as per-layer hidden states, has emerged as a promising paradigm (e.g., MemoryLLM and M+) for long-term memory in large language…

cs.CV2026

MLLMCLIP: Feature-Level Distillation of MLLM for Robust Vision-Language Representations

Jongsuk Kim, Qiyu Wu, Zhuoyuan Mao +3

Pretrained vision-language models such as CLIP excel at zero-shot recognition but often fail at compositionality, particularly attribute-object and relational structures. Recent st…

cs.SD2025

Cross-Modal Learning for Music-to-Music-Video Description Generation

Zhuoyuan Mao, Mengjie Zhao, Qiyu Wu +4

Music-to-music-video generation is a challenging task due to the intrinsic differences between the music and video modalities. The advent of powerful text-to-video diffusion models…

cs.SD2025

DeepResonance: Enhancing Multimodal Music Understanding via Music-centric Multi-way Instruction Tuning

Zhuoyuan Mao, Mengjie Zhao, Qiyu Wu +2

Recent advancements in music large language models (LLMs) have significantly improved music understanding tasks, which involve the model's ability to analyze and interpret various…

cs.SD2024★ 1 cited

OpenMU: Your Swiss Army Knife for Music Understanding

Mengjie Zhao, Zhi Zhong, Zhuoyuan Mao +5

We present OpenMU-Bench, a large-scale benchmark suite for addressing the data scarcity issue in training multimodal language models to understand music. To construct OpenMU-Bench,…

cs.CV2024★ 1 cited

GLOV: Guided Large Language Models as Implicit Optimizers for Vision Language Models

M. Jehanzeb Mirza, Mengjie Zhao, Zhuoyuan Mao +12

In this work, we propose GLOV, which enables Large Language Models (LLMs) to act as implicit optimizers for Vision-Language Models (VLMs) to enhance downstream vision tasks. GLOV p…