10 papers
ZOMP: Zeroth-Order Multi-Modal Prompt Tuning for Vision-Language Models
Sajjad Ghiasvand, Yifan Yang, Mahnoosh Alizadeh +1
Fine-tuning vision-language models such as CLIP typically requires backpropagation (BP) through the full model, which is infeasible when only forward-pass access is available, as i…
REALM: Reliable Expertise-Aware Language Model Fine-Tuning from Noisy Annotations
Sajjad Ghiasvand, Mark Beliaev, Mahnoosh Alizadeh +1
Supervised fine-tuning of large language models relies on human-annotated data, yet annotation pipelines routinely involve multiple crowdworkers of heterogeneous expertise. Standar…
MMLoP: Multi-Modal Low-Rank Prompting for Efficient Vision-Language Adaptation
Sajjad Ghiasvand, Haniyeh Ehsani Oskouie, Mahnoosh Alizadeh +1
Prompt learning has become a dominant paradigm for adapting vision-language models (VLMs) such as CLIP to downstream tasks without modifying pretrained weights. While extending pro…
Can MLLMs Critique Like Humans? Evaluating Open-Ended Aesthetic Reasoning in Multimodal Large Language Models
Sajjad Ghiasvand, Maryam Amirizaniani, Haniyeh Ehsani Oskouie +2
Open-ended aesthetic critique is a challenge for multimodal large language models (MLLMs): unlike multiple-choice aesthetic benchmarks, it has no single correct answer, and most ae…
Steady-state Based Approach to Online Non-stochastic Control
Vijeth Hebbar, Spencer Hutchinson, Mahnoosh Alizadeh +1
We study the problem of online non-stochastic control (ONC), which is the control of a linear system under adversarial disturbances and adversarial cost functions, with the aim of…
pFedMMA: Personalized Federated Fine-Tuning with Multi-Modal Adapter for Vision-Language Models
Sajjad Ghiasvand, Mahnoosh Alizadeh, Ramtin Pedarsani
Vision-Language Models (VLMs) like CLIP have demonstrated remarkable generalization in zero- and few-shot settings, but adapting them efficiently to decentralized, heterogeneous da…