5 papers
BankerToolBench: Evaluating AI Agents in End-to-End Investment Banking Workflows
Elaine Lau, Markus Dücker, Ronak Chaudhary +24
Existing AI benchmarks lack the fidelity to assess economically meaningful progress on professional workflows. To evaluate frontier AI agents in a high-value, labor-intensive profe…
Modeling Language as a Sequence of Thoughts
Nasim Borazjanizadeh, James McClelland
Transformer language models can generate strikingly natural text by modeling language as a sequence of tokens, but by relying primarily on surface-level co-occurrence statistics th…
Reliable Reasoning Beyond Natural Language
Nasim Borazjanizadeh, Steven T. Piantadosi
Despite their linguistic competence, Large Language Models (LLMs) often struggle to reason reliably and flexibly. To identify these shortcomings, we introduce the Non-Linear Reason…
Visualizing Thought: Conceptual Diagrams Enable Robust Planning in LMMs
Nasim Borazjanizadeh, Roei Herzig, Eduard Oks +3
Human reasoning relies on constructing and manipulating mental models -- simplified internal representations of situations used to understand and solve problems. Conceptual diagram…
Navigating the Labyrinth: Evaluating LLMs' Ability to Reason About Search Problems
Nasim Borazjanizadeh, Roei Herzig, Trevor Darrell +2
Large Language Models (LLMs) have recently achieved impressive performance in math and reasoning benchmarks. However, they often struggle with logic problems and puzzles that are r…