3 papers
cs.AI2026
MineCEraft: Evaluating Language Models as Construction Engineers in the World of Minecraft
Sewoong Lee, Risham Sidhu, Julia Hockenmaier +1
We introduce MineCEraft (Minecraft Construction Engineering Benchmark, pronounced mine-see-ee-raft), an easy-to-use, open-source benchmark designed to systematically evaluate the r…
cs.CL2026
Steering Instruction Hierarchies at Inference Time
Siqi Zeng, Sewoong Lee, Han Zhao +1
Instruction hierarchies are a core safety assumption of language model deployment: higher priority inputs, such as system prompts, should override conflicting lower priority inputs…
cs.LG2025
Evaluating and Designing Sparse Autoencoders by Approximating Quasi-Orthogonality
Sewoong Lee, Adam Davies, Marc E. Canby +1
Sparse autoencoders (SAEs) are widely used in mechanistic interpretability research for large language models; however, the state-of-the-art method of using -sparse autoencoders…