9 citations · 23 across the 17 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
MineCEraft: Evaluating Language Models as Construction Engineers in the World of Minecraft
Sewoong Lee, Risham Sidhu, Julia Hockenmaier +1
We introduce MineCEraft (Minecraft Construction Engineering Benchmark, pronounced mine-see-ee-raft), an easy-to-use, open-source benchmark designed to systematically evaluate the r…
cs.AI2025
Evaluating Legal Reasoning Traces with Legal Issue Tree Rubrics
Jinu Lee, Kyoung-Woon On, Simeng Han +2
Evaluating the quality of LLM-generated reasoning traces in expert domains (e.g., law) is essential for ensuring credibility and explainability, yet remains challenging due to the…