3 citations · 3 across the 2 of their papers we have counts for
2 papers
cs.CL2023
CLEAN-EVAL: Clean Evaluation on Contaminated Large Language Models
Wenhong Zhu, Hongkun Hao, Zhiwei He +6
We are currently in an era of fierce competition among various large language models (LLMs) continuously pushing the boundaries of benchmark performance. However, genuinely assessi…
cs.AI2023★ 3 cited
TPE: Towards Better Compositional Reasoning over Conceptual Tools with Multi-persona Collaboration
Hongru Wang, Huimin Wang, Lingzhi Wang +6
Large language models (LLMs) have demonstrated exceptional performance in planning the use of various functional tools, such as calculators and retrievers, particularly in question…