3 papers
cs.CV2025
KARST: Multi-Kernel Kronecker Adaptation with Re-Scaling Transmission for Visual Classification
Yue Zhu, Haiwen Diao, Shang Gao +2
Fine-tuning pre-trained vision models for specific tasks is a common practice in computer vision. However, this process becomes more expensive as models grow larger. Recently, para…
cs.AI2024
LLMs Can Evolve Continually on Modality for X-Modal Reasoning
Jiazuo Yu, Haomiao Xiong, Lu Zhang +7
Multimodal Large Language Models (MLLMs) have gained significant attention due to their impressive capabilities in multimodal understanding. However, existing methods rely heavily…
cs.CV2024
GSSF: Generalized Structural Sparse Function for Deep Cross-modal Metric Learning
Haiwen Diao, Ying Zhang, Shang Gao +3
Cross-modal metric learning is a prominent research topic that bridges the semantic heterogeneity between vision and language. Existing methods frequently utilize simple cosine or…