6 citations · 6 across the 2 of their papers we have counts for
3 papers
cs.CL2026
BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks
Nishant Balepur, Bhavya Rajasekaran, Jane Oh +7
Multiple-choice question answering (MCQA) is standard in NLP, but benchmarks lack rigorous quality control. We present BenchMarker, an education-inspired toolkit using LLM judges t…
cs.AI2024★ 6 cited
Can AI Be as Creative as Humans?
Haonan Wang, James Zou, Michael Mozer +8
Creativity serves as a cornerstone for societal progress and innovation. With the rise of advanced generative AI models capable of tasks once reserved for human creativity, the stu…
cs.LG2023
Prompt Optimization via Adversarial In-Context Learning
Xuan Long Do, Yiran Zhao, Hannah Brown +6
We propose a new method, Adversarial In-Context Learning (adv-ICL), to optimize prompt for in-context learning (ICL) by employing one LLM as a generator, another as a discriminator…