3 papers
cs.LG2026
ExtractBench: A Benchmark and Evaluation Methodology for Complex Structured Extraction
Nick Ferguson, Josh Pennington, Narek Beghian +4
Unstructured documents like PDFs contain valuable structured information, but downstream systems require this data in reliable, standardized formats. LLMs are increasingly deployed…
cs.CL2026
Exploring the Meta-level Reasoning of Large Language Models via a Tool-based Multi-hop Tabular Question Answering Task
Nick Ferguson, Alan Bundy, Kwabena Nuamah
Recent advancements in Large Language Models (LLMs) are increasingly focused on "reasoning" ability, a concept with many overlapping definitions in the LLM discourse. We take a mor…
cs.CL2025
Evaluating the Meta- and Object-Level Reasoning of Large Language Models for Question Answering
Nick Ferguson, Liane Guillou, Alan Bundy +1
Large Language Models (LLMs) excel in natural language tasks but still face challenges in Question Answering (QA) tasks requiring complex, multi-step reasoning. We outline the type…