1 paper · 1 filter
Aihua Pei, Zehua Yang, Shunan Zhu +3
Existing frameworks for assessing robustness of large language models (LLMs) overly depend on specific benchmarks, increasing costs and failing to evaluate performance of LLMs in p…