3 papers
cs.LG2026
Fork-Think with Confidence
Zena Al-Khalili, Rafi Hakim, Dietrich Klakow +1
Parallel thinking has enjoyed great success for boosting LLM performance on reasoning tasks without the need for any re-training. However, existing methods follow a think-first-the…
cs.AI2025
PricingLogic: Evaluating LLMs Reasoning on Complex Tourism Pricing Tasks
Yunuo Liu, Dawei Zhu, Zena Al-Khalili +5
We present PricingLogic, the first benchmark that probes whether Large Language Models(LLMs) can reliably automate tourism-related prices when multiple, overlapping fare rules appl…
cs.CL2025
Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics
Zena Al-Khalili, Nick Howell, Dietrich Klakow
Assisting LLMs with code generation improved their performance on mathematical reasoning tasks. However, the evaluation of code-assisted LLMs is generally restricted to execution c…