13 citations · 13 across the 2 of their papers we have counts for
2 papers
cs.CV2024
Image2Struct: Benchmarking Structure Extraction for Vision-Language Models
Josselin Somerville Roberts, Tony Lee, Chi Heem Wong +3
We introduce Image2Struct, a benchmark to evaluate vision-language models (VLMs) on extracting structure from images. Our benchmark 1) captures real-world use cases, 2) is fully au…
cs.CV2023★ 13 cited
Holistic Evaluation of Text-To-Image Models
Tony Lee, Michihiro Yasunaga, Chenlin Meng +15
The stunning qualitative improvement of recent text-to-image models has led to their widespread attention and adoption. However, we lack a comprehensive quantitative understanding…