1 paper
Meilong Xu, Di Fu, Jiaxing Zhang +7
Vision Language Models (VLMs) are becoming increasingly integral to multimedia understanding; however, they often struggle with domain-specific video classification tasks, particul…