219 citations · 227 across the 7 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
Nevermind: Instruction Override and Moderation in Large Language Models
Edward Kim
Given the impressive capabilities of recent Large Language Models (LLMs), we investigate and benchmark the most popular proprietary and different sized open source models on the ta…
cs.CL2023★ 4 cited
Which Prompts Make The Difference? Data Prioritization For Efficient Human LLM Evaluation
Meriem Boubdir, Edward Kim, Beyza Ermis +2
Human evaluation is increasingly critical for assessing large language models, capturing linguistic nuances, and reflecting user preferences more accurately than traditional automa…