collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2025

Synthetic Dataset for Evaluating Complex Compositional Knowledge for Natural Language Inference

Sushma Anand Akoju, Robert Vacareanu, Haris Riaz +2

We introduce a synthetic dataset called Sentences Involving Complex Compositional Knowledge (SICCK) and a novel analysis that investigates the performance of Natural Language Infer…

cs.CL2025

Data Contamination Quiz: A Tool to Detect and Estimate Contamination in Large Language Models

Shahriar Golchin, Mihai Surdeanu

We propose the Data Contamination Quiz (DCQ), a simple and effective approach to detect data contamination in large language models (LLMs) and estimate the amount of it. Specifical…

cs.CL2025

Memorization in In-Context Learning

Shahriar Golchin, Mihai Surdeanu, Steven Bethard +2

In-context learning (ICL) has proven to be an effective strategy for improving the performance of large language models (LLMs) with no additional training. However, the exact mecha…

cs.CL2024

From Words to Numbers: Your Large Language Model Is Secretly A Capable Regressor When Given In-Context Examples

Robert Vacareanu, Vlad-Andrei Negru, Vasile Suciu +1

We analyze how well pre-trained large language models (e.g., Llama2, GPT-4, Claude 3, etc) can do linear and non-linear regression when given in-context examples, without any addit…

cs.CL2024

Data Contamination Report from the 2024 CONDA Shared Task

Oscar Sainz, Iker García-Ferrero, Alon Jacovi +25

The 1st Workshop on Data Contamination (CONDA 2024) focuses on all relevant aspects of data contamination in natural language processing, where data contamination is understood as…