1 paper
Batuhan Yeltekin, Daniel Bauer
We evaluate 13 LLMs on generating, solving, simulating, and diagnosing novice programming misconceptions on code-tracing problems. While frontier models reliably perform these task…