From the 1 of 7 linked papers with an AI index.
7 papers
Video to All-in-focus Image Reconstruction Algorithm for Automated Microscopic Urinalysis
Chinmay Nema, Hari Om Aggrawal, Dipam Goswami +2
The paper presents a method that records a short video while manually adjusting focus and reconstructs an all‑in‑focus image from the frames, enabling automated deep‑learning based…
IsoCLIP: Decomposing CLIP Projectors for Efficient Intra-modal Alignment
Simone Magistri, Dipam Goswami, Marco Mistretta +3
Vision-Language Models like CLIP are extensively used for inter-modal tasks which involve both visual and text modalities. However, when the individual modality encoders are applie…
Training data attribution in diffusion models via mirrored unlearning and noise-consistent skew
Joan SerrÃ, Dipam Goswami, Fabio Morreale +2
Training data attribution (TDA) should enable generative model interpretability and foster a variety of related downstream tasks. Nonetheless, current TDA approaches lack reliabili…
Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting
Yuyang Liu, Qiuhe Hong, Linlan Huang +6
Vision-language models (VLMs), spanning predictive architectures to generative Multimodal Large Language Models (MLLMs), have revolutionized artificial intelligence through powerfu…
Cross-Modal Prototype Alignment and Mixing for Training-Free Few-Shot Classification
Dipam Goswami, Simone Magistri, Gido M. van de Ven +4
Vision-language models (VLMs) like CLIP are trained with the objective of aligning text and image pairs. To improve CLIP-based few-shot image classification, recent works have obse…
Covariances for Free: Exploiting Mean Distributions for Training-free Federated Learning
Dipam Goswami, Simone Magistri, Kai Wang +3
Using pre-trained models has been found to reduce the effect of data heterogeneity and speed up federated learning algorithms. Recent works have explored training-free methods usin…