1 citations · 1 across the 4 of their papers we have counts for
7 papers · 1 filter
Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation
Yue Yang, Ajay Patel, Matt Deitke +8
Reasoning about images with rich text, such as charts and documents, is a critical application of vision-language models (VLMs). However, VLMs often struggle in these domains due t…
ViUniT: Visual Unit Tests for More Robust Visual Programming
Artemis Panagopoulou, Honglu Zhou, Silvio Savarese +4
Programming based approaches to reasoning tasks have substantially expanded the types of questions models can answer about visual scenes. Yet on benchmark visual reasoning data, wh…
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Matt Deitke, Christopher Clark, Sangho Lee +47
Today's most advanced vision-language models (VLMs) remain proprietary. The strongest open-weight models rely heavily on synthetic data from proprietary VLMs to achieve good perfor…
A Textbook Remedy for Domain Shifts: Knowledge Priors for Medical Image Analysis
Yue Yang, Mona Gandhi, Yufei Wang +5
While deep networks have achieved broad success in analyzing natural images, when applied to medical scans, they often fail in unexcepted situations. We investigate this challenge…
CoMo: Controllable Motion Generation through Language Guided Pose Code Editing
Yiming Huang, Weilin Wan, Yue Yang +3
Text-to-motion models excel at efficient human motion generation, but existing approaches lack fine-grained controllability over the generation process. Consequently, modifying sub…
LLM-based Hierarchical Concept Decomposition for Interpretable Fine-Grained Image Classification
Renyi Qu, Mark Yatskar
(Renyi Qu's Master's Thesis) Recent advancements in interpretable models for vision-language tasks have achieved competitive performance; however, their interpretability often suff…