5 papers
MAOAM: Unified Object and Material Selection with Vision-Language Models
Jaden Park, Valentin Deschaintre, Jason Kuen +5
Selection is a core operation in interactive image editing. To be practical, a user should be able to specify and disambiguate the desired selection region through either text or c…
Exploration and Exploitation Errors Are Measurable for Language Model Agents
Jaden Park, Jungtaek Kim, Jongwon Jeong +3
Language Model (LM) agents are increasingly used in complex open-ended decision-making tasks, from AI coding to physical AI. A core requirement in these settings is the ability to…
Contamination Detection for VLMs using Multi-Modal Semantic Perturbation
Jaden Park, Mu Cai, Feng Yao +3
Recent advances in Vision-Language Models (VLMs) have achieved state-of-the-art performance on numerous benchmark tasks. However, the use of internet-scale, often proprietary, pret…
Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models
Hyunsik Chae, Seungwoo Yoon, Jaden Park +5
Recent Vision-Language Models (VLMs) have demonstrated impressive multimodal comprehension and reasoning capabilities, yet they often struggle with trivially simple visual tasks. I…
TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
Mu Cai, Reuben Tan, Jianrui Zhang +12
Understanding fine-grained temporal dynamics is crucial for multimodal video comprehension and generation. Due to the lack of fine-grained temporal annotations, existing video benc…