1 paper
Xin Tang, Youfang Han, Fangfei Gou +7
Vision-Language Models (VLMs) excel in diverse multimodal tasks. However, user requirements vary across scenarios, which can be categorized into fast response, high-quality output,…