10 papers
VisGym: Diverse, Customizable, Scalable Environments for Multimodal Agents
Zirui Wang, Junyi Zhang, Jiaxin Ge +9
Modern Vision-Language Models (VLMs) remain poorly characterized in multi-step visual interactions, particularly in how they integrate perception, memory, and action over long hori…
Constantly Improving Image Models Need Constantly Improving Benchmarks
Jiaxin Ge, Grace Luo, Heekyung Lee +7
Recent advances in image generation, often driven by proprietary systems like GPT-4o Image Gen, regularly introduce new capabilities that reshape how users interact with these mode…
Do What? Teaching Vision-Language-Action Models to Reject the Impossible
Wen-Han Hsieh, Elvis Hsieh, Dantong Niu +3
Recently, Vision-Language-Action (VLA) models have demonstrated strong performance on a range of robotic tasks. These models rely on multimodal inputs, with language instructions p…
LISAT: Language-Instructed Segmentation Assistant for Satellite Imagery
Jerome Quenum, Wen-Han Hsieh, Tsung-Han Wu +3
Segmentation models can recognize a pre-defined set of objects in images. However, models that can reason over complex user queries that implicitly refer to multiple objects of int…
Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint
Heekyung Lee, Jiaxin Ge, Tsung-Han Wu +3
Rebus puzzles, visual riddles that encode language through imagery, spatial arrangement, and symbolic substitution, pose a unique challenge to current vision-language models (VLMs)…
REOrdering Patches Improves Vision Models
Declan Kutscher, David M. Chan, Yutong Bai +2
Sequence models such as transformers require inputs to be represented as one-dimensional sequences. In vision, this typically involves flattening images using a fixed row-major (ra…