2 papers
cs.CL2026
SciTaRC: A Plan-Annotated Scientific Tabular QA Benchmark for Language Reasoning and Complex Computation
Hexuan Wang, Yaxuan Ren, Srikar Bommireddypalli +5
We introduce SciTaRC, an expert-authored benchmark for question answering over scientific tables that targets composite, multi-step reasoning. To enable fine-grained diagnostic ana…
astro-ph.IM2024
Designing an Evaluation Framework for Large Language Models in Astronomy Research
John F. Wu, Alina Hyk, Kiera McCormick +15
Large Language Models (LLMs) are shifting how scientific research is done. It is imperative to understand how researchers interact with these models and how scientific sub-communit…