562 citations · 563 across the 2 of their papers we have counts for
2 papers
cs.CV2023★ 1 cited
Simple Token-Level Confidence Improves Caption Correctness
Suzanne Petryk, Spencer Whitehead, Joseph E. Gonzalez +3
The ability to judge whether a caption correctly describes an image is a critical part of vision-language understanding. However, state-of-the-art models often misinterpret the cor…
cs.CV2023★ 562 cited
Segment Anything
Alexander Kirillov, Eric Mintun, Nikhila Ravi +9
We introduce the Segment Anything (SA) project: a new task, model, and dataset for image segmentation. Using our efficient model in a data collection loop, we built the largest seg…