10 citations · 21 across the 4 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2024★ 10 cited
Using Counterfactual Tasks to Evaluate the Generality of Analogical Reasoning in Large Language Models
Martha Lewis, Melanie Mitchell
Large language models (LLMs) have performed well on several reasoning benchmarks, including ones that test analogical reasoning abilities. However, it has been debated whether they…
cs.AI2022★ 7 cited
Evaluating Understanding on Conceptual Abstraction Benchmarks
Victor Vikram Odouard, Melanie Mitchell
A long-held objective in AI is to build systems that understand concepts in a humanlike way. Setting aside the difficulty of building such a system, even trying to evaluate one is…