37 citations · 51 across the 10 of their papers we have counts for
11 papers
The Impact of Post-training on Data Contamination
Muhammed Yusuf Kocyigit, Caglar Yildirim
We present a controlled study of how dataset contamination interacts with the post-training stages now standard in large language model training pipelines. Starting from clean chec…
Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation
Muhammed Yusuf Kocyigit, Eleftheria Briakou, Daniel Deutsch +3
Data contamination -- the accidental consumption of evaluation examples within the pre-training data -- can undermine the validity of evaluation benchmarks. In this paper, we prese…
Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?
Aaditya K. Singh, Muhammed Yusuf Kocyigit, Andrew Poulton +4
Hampering the interpretation of benchmark scores, evaluation data contamination has become a growing concern in the evaluation of LLMs, and an active area of research studies its e…
Western, Religious or Spiritual: An Evaluation of Moral Justification in Large Language Models
Eyup Engin Kucuk, Muhammed Yusuf Kocyigit
The increasing success of Large Language Models (LLMs) in variety of tasks lead to their widespread use in our lives which necessitates the examination of these models from differe…
A Novel Method for Analysing Racial Bias: Collection of Person Level References
Muhammed Yusuf Kocyigit, Anietie Andy, Derry Wijaya
Long term exposure to biased content in literature or media can significantly influence people's perceptions of reality, leading to the development of implicit biases that are diff…
AugCSE: Contrastive Sentence Embedding with Diverse Augmentations
Zilu Tang, Muhammed Yusuf Kocyigit, Derry Wijaya
Data augmentation techniques have been proven useful in many applications in NLP fields. Most augmentations are task-specific, and cannot be used as a general-purpose tool. In our…