2 citations · 5 across the 17 of their papers we have counts for
1 paper · 1 filter
JiaRu Wu, Mingwei Liu
Large language models (LLMs) have shown remarkable performance on various tasks, but existing evaluation benchmarks are often static and insufficient to fully assess their robustne…