4 papers
Failure Modes in Multi-Hop QA: The Weakest Link Effect and the Recognition Bottleneck
Meiru Zhang, Zaiqiao Meng, Nigel Collier
Despite scaling to massive context windows, Large Language Models (LLMs) struggle with multi-hop reasoning due to inherent position bias, which causes them to overlook information…
DRIFT: Decompose, Retrieve, Illustrate, then Formalize Theorems
Meiru Zhang, Philipp Borchert, Milan Gritta +1
Automating the formalization of mathematical statements for theorem proving remains a major challenge for Large Language Models (LLMs). LLMs struggle to identify and utilize the pr…
A Benchmark for Deep Information Synthesis
Debjit Paul, Daniel Murphy, Milan Gritta +14
Large language model (LLM)-based agents are increasingly used to solve complex tasks involving tool use, such as web browsing, code execution, and data analysis. However, current e…
Breaking Thought Patterns: A Multi-Dimensional Reasoning Framework for LLMs
Xintong Tang, Meiru Zhang, Shang Xiao +5
Large language models (LLMs) are often constrained by rigid reasoning processes, limiting their ability to generate creative and diverse responses. To address this, a novel framewo…