1 citations · 1 across the 1 of their papers we have counts for
3 papers
cs.LG2025
An Empirical Study of Federated Prompt Learning for Vision Language Model
Zhihao Wang, Wenke Huang, Tian Chen +7
The Vision Language Model (VLM) excels in aligning vision and language representations, and prompt learning has emerged as a key technique for adapting such models to downstream ta…
cs.CL2025
Keeping Yourself is Important in Downstream Tuning Multimodal Large Language Model
Wenke Huang, Jian Liang, Xianda Guo +14
Multi-modal Large Language Models (MLLMs) integrate visual and linguistic reasoning to address complex tasks such as image captioning and visual question answering. While MLLMs dem…
cs.CL2024★ 1 cited
Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning
Wenke Huang, Jian Liang, Zekun Shi +6
Multimodal Large Language Model (MLLM) have demonstrated strong generalization capabilities across diverse distributions and tasks, largely due to extensive pre-training datasets.…