1 paper · 1 filter
Linying Yang, Vik Shirvaikar, Oscar Clivio +1
Numerous benchmarks aim to evaluate the capabilities of Large Language Models (LLMs) for causal inference and reasoning. However, many of them can likely be solved through the retr…