2 papers
cs.IR2025
Embedding the Teacher: Distilling vLLM Preferences for Scalable Image Retrieval
Eric He, Akash Gupta, Adian Liusie +4
Text--image retrieval is necessary for applications such as product recommendation. Embedding-based approaches like CLIP enable efficient large-scale retrieval via vector similarit…
cs.CL2025
Probing the Limits of Stylistic Alignment in Vision-Language Models
Asma Farajidizaji, Akash Gupta, Vatsal Raina
Vision-language models are increasingly used to generate image captions in specific styles, such as humor or romantic. However, these transformer-based models often struggle with t…