3 papers
cs.CL2026
BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks
Nishant Balepur, Bhavya Rajasekaran, Jane Oh +7
Multiple-choice question answering (MCQA) is standard in NLP, but benchmarks lack rigorous quality control. We present BenchMarker, an education-inspired toolkit using LLM judges t…
cs.CL2026
Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers
Nishant Balepur, Atrey Desai, Rachel Rudinger
Large language models (LLMs) now give reasoning before answering, excelling in tasks like multiple-choice question answering (MCQA). Yet, a concern is that LLMs do not solve MCQs a…
cs.CL2026
Filling in the Mechanisms: How do LMs Learn Filler-Gap Dependencies under Developmental Constraints?
Atrey Desai, Sathvik Nair
For humans, filler-gap dependencies require a shared representation across different syntactic constructions. Although causal analyses suggest this may also be true for LLMs (Bogur…