2 papers
cs.CV2024
Image2Struct: Benchmarking Structure Extraction for Vision-Language Models
Josselin Somerville Roberts, Tony Lee, Chi Heem Wong +3
We introduce Image2Struct, a benchmark to evaluate vision-language models (VLMs) on extracting structure from images. Our benchmark 1) captures real-world use cases, 2) is fully au…
cs.CV2023
A Skeleton-based Approach For Rock Crack Detection Towards A Climbing Robot Application
Josselin Somerville Roberts, Paul-Emile Giacomelli, Yoni Gozlan +1
Conventional wheeled robots are unable to traverse scientifically interesting, but dangerous, cave environments. Multi-limbed climbing robot designs, such as ReachBot, are able to…