9 papers
Personal AI Agent for Camera Roll VQA
Thao Nguyen, Krishna Kumar Singh, Donghyun Kim +2
We study the personal camera roll visual question answering setting. In this setting, a conversational AI assistant can access a user's personal camera roll and retrieve relevant p…
MAOAM: Unified Object and Material Selection with Vision-Language Models
Jaden Park, Valentin Deschaintre, Jason Kuen +5
Selection is a core operation in interactive image editing. To be practical, a user should be able to specify and disambiguate the desired selection region through either text or c…
MuRF: Unlocking the Multi-Scale Potential of Vision Foundation Models
Bocheng Zou, Mu Cai, Mark Stanley +2
Vision Foundation Models (VFMs) have become the cornerstone of modern computer vision, offering robust representations across a wide array of tasks. While recent advances allow the…
Contamination Detection for VLMs using Multi-Modal Semantic Perturbation
Jaden Park, Mu Cai, Feng Yao +3
Recent advances in Vision-Language Models (VLMs) have achieved state-of-the-art performance on numerous benchmark tasks. However, the use of internet-scale, often proprietary, pret…
CHARTOM: A Visual Theory-of-Mind Benchmark for LLMs on Misleading Charts
Shubham Bharti, Shiyun Cheng, Jihyun Rho +5
We introduce CHARTOM, a visual theory-of-mind benchmark designed to evaluate multimodal large language models' capability to understand and reason about misleading data visualizati…
Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models
Hyunsik Chae, Seungwoo Yoon, Jaden Park +5
Recent Vision-Language Models (VLMs) have demonstrated impressive multimodal comprehension and reasoning capabilities, yet they often struggle with trivially simple visual tasks. I…