5 papers
Reassessing Extractive QA Datasets at Scale: LLM-as-a-Judge and In-Depth Analyses
Xanh Ho, Jiahao Huang, Florian Boudin +1
Extractive QA tasks are commonly evaluated using Exact Match (EM) and F1-score, but these metrics often fail to reflect true model performance. Recent studies have proposed using l…
Tailoring the Curriculum: Student-Centered Reasoning Distillation via Dynamic Data-Model Compatibility
Jiahao Huang, Fei Cheng, Junfeng Jiang +1
Reasoning distillation transfers complex reasoning abilities from large language models (LLMs) to smaller ones, yet its success depends on how well the training data align with the…
BenchTrace: A Benchmark for Testing Reflection Ability and Controlled Evolution in LLM Agents
Jiahao Huang, Fei Cheng, Junfeng Jiang +2
Self-evolving agents improve over time by reflecting on past failures, but existing evaluation is limited in two ways: it measures only task scores, leaving reflection quality unkn…
Med-CoReasoner: Reducing Language Disparities in Medical Reasoning via Language-Informed Co-Reasoning
Fan Gao, Sherry T. Tong, Jiwoong Sohn +11
While reasoning-enhanced large language models perform strongly on English medical tasks, a persistent multilingual gap remains, with substantially weaker reasoning in local langua…
TactfulToM: Do LLMs Have the Theory of Mind Ability to Understand White Lies?
Yiwei Liu, Emma Jane Pretty, Jiahao Huang +1
While recent studies explore Large Language Models' (LLMs) performance on Theory of Mind (ToM) reasoning tasks, research on ToM abilities that require more nuanced social context i…