most citedmodeLing: A Novel Dataset for Testing Linguistic Reasoning in Language Models

2 citations · 2 across the 1 of their papers we have counts for

collaborators

5 papers

cs.AI2024

P-FOLIO: Evaluating and Improving Logical Reasoning with Abundant Human-Written Reasoning Chains

Simeng Han, Aaron Yu, Rui Shen +13

Existing methods on understanding the capabilities of LLMs in logical reasoning rely on binary entailment classification or synthetically derived rationales, which are not sufficie…

cs.CL20242 cited

modeLing: A Novel Dataset for Testing Linguistic Reasoning in Language Models

Nathan A. Chi, Teodor Malchev, Riley Kong +5

We introduce modeLing, a novel benchmark of Linguistics Olympiad-style puzzles which tests few-shot reasoning in AI systems. Solving these puzzles necessitates inferring aspects of…

cs.CL2023

Ascle: A Python Natural Language Processing Toolkit for Medical Text Generation

Rui Yang, Qingcheng Zeng, Keen You +13

This study introduces Ascle, a pioneering natural language processing (NLP) toolkit designed for medical text generation. Ascle is tailored for biomedical researchers and healthcar…

cs.CL2023

Benchmarking Generation and Evaluation Capabilities of Large Language Models for Instruction Controllable Summarization

Yixin Liu, Alexander R. Fabbri, Jiawen Chen +7

While large language models (LLMs) can already achieve strong performance on standard generic summarization benchmarks, their performance on more complex summarization task setting…

cs.CL2023

Fair Abstractive Summarization of Diverse Perspectives

Yusen Zhang, Nan Zhang, Yixin Liu +9

People from different social and demographic groups express diverse perspectives and conflicting opinions on a broad set of topics such as product reviews, healthcare, law, and pol…