2 citations · 5 across the 8 of their papers we have counts for
3 papers · 1 filter
Image2Struct: Benchmarking Structure Extraction for Vision-Language Models
Josselin Somerville Roberts, Tony Lee, Chi Heem Wong +3
We introduce Image2Struct, a benchmark to evaluate vision-language models (VLMs) on extracting structure from images. Our benchmark 1) captures real-world use cases, 2) is fully au…
VHELM: A Holistic Evaluation of Vision Language Models
Tony Lee, Haoqin Tu, Chi Heem Wong +8
Current benchmarks for assessing vision-language models (VLMs) often focus on their perception or problem-solving capabilities and neglect other critical aspects such as fairness,…
A Skeleton-based Approach For Rock Crack Detection Towards A Climbing Robot Application
Josselin Somerville Roberts, Paul-Emile Giacomelli, Yoni Gozlan +1
Conventional wheeled robots are unable to traverse scientifically interesting, but dangerous, cave environments. Multi-limbed climbing robot designs, such as ReachBot, are able to…