3 papers
cs.AI2024
P-FOLIO: Evaluating and Improving Logical Reasoning with Abundant Human-Written Reasoning Chains
Simeng Han, Aaron Yu, Rui Shen +13
Existing methods on understanding the capabilities of LLMs in logical reasoning rely on binary entailment classification or synthetically derived rationales, which are not sufficie…
cs.CL2024
Benchmarking Generation and Evaluation Capabilities of Large Language Models for Instruction Controllable Summarization
Yixin Liu, Alexander R. Fabbri, Jiawen Chen +7
While large language models (LLMs) can already achieve strong performance on standard generic summarization benchmarks, their performance on more complex summarization task setting…
cs.CL2024
modeLing: A Novel Dataset for Testing Linguistic Reasoning in Language Models
Nathan A. Chi, Teodor Malchev, Riley Kong +5
We introduce modeLing, a novel benchmark of Linguistics Olympiad-style puzzles which tests few-shot reasoning in AI systems. Solving these puzzles necessitates inferring aspects of…