4 papers
MeDxAgent: Multi-Agent Consultation for Interactive Medical Diagnosis
Akshat Sanghvi, Naren Akash, Raza Imam +2
Large language models (LLMs) are increasingly used for health-related decision support. Yet most evaluations treat diagnosis as a single-shot task with complete information provide…
Better Think Thrice: Learning to Reason Causally with Double Counterfactual Consistency
Victoria Lin, Xinnuo Xu, Rachel Lawrence +4
Despite their strong performance on reasoning benchmarks, large language models (LLMs) have proven brittle when presented with counterfactual questions, suggesting weaknesses in th…
RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation
Xinnuo Xu, Rachel Lawrence, Kshitij Dubey +7
Recent Large Language Models (LLMs) have reported high accuracy on reasoning benchmarks. However, it is still unclear whether the observed results arise from true reasoning or from…
DeduCE: Deductive Consistency as a Framework to Evaluate LLM Reasoning
Atharva Pandey, Kshitij Dubey, Rahul Sharma +1
Despite great performance on Olympiad-level reasoning problems, frontier large language models can still struggle on high school math when presented with novel problems outside sta…