2 papers
cs.CL2026
MUDDLE: Measuring Understanding of Documents under Distractor and Length Effects
Jason Luo, Saibilila Abudukelimu, Judy Song +4
Document question-answering systems increasingly answer questions over collections of retrieved documents rather than one clean source, so robustness to distracting context matters…
cs.CL2026
NL2AGBench: Benchmarking LLM Auto-Formalization for AlphaGeometry
Samuel Xiao, Judy Song, Rory Hu +1
Recent advances in large language models (LLMs) have demonstrated strong capabilities in natural language understanding and mathematical reasoning. However, their ability to transl…