7 papers
SAMPLe: SAM-based Optimizer for Prompt Learning in VLMs
Hossein Rajoli, Fatemeh Lotfi, Niloufar Alipour Talemi +3
Pre-trained Vision-Language Models (VLMs) like CLIP have proven highly effective as foundation models for various downstream applications. However, prompt learning in VLMs encounte…
Hyper-ICL: Attention Calibration with Hyperbolic Anchor Distillation for Multimodal In-Context Learning
Niloufar Alipour Talemi, Hossein Kashiani, Fatemeh Afghah
Multimodal In-Context Learning (ICL) has emerged as a practical inference paradigm for Multimodal Large Language Models, where a small set of interleaved image-text In-Context Demo…
PromptMAD: Cross-Modal Prompting for Multi-Class Visual Anomaly Localization
Duncan McCain, Hossein Kashiani, Fatemeh Afghah
Visual anomaly detection in multi-class settings poses significant challenges due to the diversity of object categories, the scarcity of anomalous examples, and the presence of cam…
FreqDebias: Towards Generalizable Deepfake Detection via Consistency-Driven Frequency Debiasing
Hossein Kashiani, Niloufar Alipour Talemi, Fatemeh Afghah
Deepfake detectors often struggle to generalize to novel forgery types due to biases learned from limited training data. In this paper, we identify a new type of model bias in the…
DiSa: Directional Saliency-Aware Prompt Learning for Generalizable Vision-Language Models
Niloufar Alipour Talemi, Hossein Kashiani, Hossein R. Nowdeh +1
Prompt learning has emerged as a powerful paradigm for adapting vision-language models such as CLIP to downstream tasks. However, existing methods often overfit to seen data, leadi…
ROADS: Robust Prompt-driven Multi-Class Anomaly Detection under Domain Shift
Hossein Kashiani, Niloufar Alipour Talemi, Fatemeh Afghah
Recent advancements in anomaly detection have shifted focus towards Multi-class Unified Anomaly Detection (MUAD), offering more scalable and practical alternatives compared to trad…