Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
MineCEraft: Evaluating Language Models as Construction Engineers in the World of Minecraft
Sewoong Lee, Risham Sidhu, Julia Hockenmaier +1
We introduce MineCEraft (Minecraft Construction Engineering Benchmark, pronounced mine-see-ee-raft), an easy-to-use, open-source benchmark designed to systematically evaluate the r…
cs.AI2026
DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings
Yoonhwa Jung, Junryu Fu, Mani Golparvar-Fard
We introduce DrawingVQA, the first benchmark designed to evaluate multimodal large language models (MLLMs) on real-world construction drawings -- a core media in architecture, civi…