27 citations · 48 across the 20 of their papers we have counts for
4 papers · 1 filter
Elo Uncovered: Robustness and Best Practices in Language Model Evaluation
Meriem Boubdir, Edward Kim, Beyza Ermis +2
In Natural Language Processing (NLP), the Elo rating system, originally designed for ranking players in dynamic games such as chess, is increasingly being used to evaluate Large La…
Which Prompts Make The Difference? Data Prioritization For Efficient Human LLM Evaluation
Meriem Boubdir, Edward Kim, Beyza Ermis +2
Human evaluation is increasingly critical for assessing large language models, capturing linguistic nuances, and reflecting user preferences more accurately than traditional automa…
Goodtriever: Adaptive Toxicity Mitigation with Retrieval-augmented Models
Luiza Pozzobon, Beyza Ermis, Patrick Lewis +1
Considerable effort has been dedicated to mitigating toxicity, but existing methods often require drastic modifications to model parameters or the use of computationally intensive…
On the Challenges of Using Black-Box APIs for Toxicity Evaluation in Research
Luiza Pozzobon, Beyza Ermis, Patrick Lewis +1
Perception of toxicity evolves over time and often differs between geographies and cultural backgrounds. Similarly, black-box commercially available APIs for detecting toxicity, su…