Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles
Eric Slyman, Mehrab Tanjim, Kushal Kafle +1
Multimodal large language models (MLLMs) are increasingly used to evaluate text-to-image (TTI) generation systems, providing automated judgments based on visual and textual context…
cs.CV2024
Hijacking Vision-and-Language Navigation Agents with Adversarial Environmental Attacks
Zijiao Yang, Xiangxi Shi, Eric Slyman +1
Assistive embodied agents that can be instructed in natural language to perform tasks in open-world environments have the potential to significantly impact labor tasks like manufac…
cs.CV2024
You Never Know: Quantization Induces Inconsistent Biases in Vision-Language Foundation Models
Eric Slyman, Anirudh Kanneganti, Sanghyun Hong +1
We study the impact of a standard practice in compressing foundation vision-language models - quantization - on the models' ability to produce socially-fair outputs. In contrast to…