106 citations · 107 across the 2 of their papers we have counts for
2 papers
cs.CL2024★ 1 cited
WTU-EVAL: A Whether-or-Not Tool Usage Evaluation Benchmark for Large Language Models
Kangyun Ning, Yisong Su, Xueqiang Lv +4
Although Large Language Models (LLMs) excel in NLP tasks, they still need external tools to extend their ability. Current research on tool learning with LLMs often assumes mandator…
cs.CL2023★ 106 cited
Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4
Hanmeng Liu, Ruoxi Ning, Zhiyang Teng +3
Harnessing logical reasoning ability is a comprehensive natural language understanding endeavor. With the release of Generative Pretrained Transformer 4 (GPT-4), highlighted as "ad…