4 papers
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Jonathan Roberts, Mohammad Reza Taesiri, Ansh Sharma +31
Large Multimodal Models (LMMs) exhibit shortfalls when interpreting images and, by some measures, have poorer spatial cognition than young children or animals. Despite this, they a…
Soft Contextualized Encoder For User Defined Text Classification
Charu Maheshwari, Vyas Raina
User-Defined Text Classification (UDTC) considers the challenge of classifying input text to user-specified, previously unseen classes, a setting that arises frequently in real-wor…
Embedding the Teacher: Distilling vLLM Preferences for Scalable Image Retrieval
Eric He, Akash Gupta, Adian Liusie +4
Text--image retrieval is necessary for applications such as product recommendation. Embedding-based approaches like CLIP enable efficient large-scale retrieval via vector similarit…
Universal Acoustic Adversarial Attacks for Flexible Control of Speech-LLMs
Rao Ma, Mengjie Qian, Vyas Raina +2
The combination of pre-trained speech encoders with large language models has enabled the development of speech LLMs that can handle a wide range of spoken language processing task…