21 citations · 28 across the 5 of their papers we have counts for
5 papers
CIBench: Evaluating Your LLMs with a Code Interpreter Plugin
Chuyu Zhang, Songyang Zhang, Yingfan Hu +8
While LLM-Based agents, which use external tools to solve complex problems, have made significant progress, benchmarking their ability is challenging, thereby hindering a clear und…
InternLM-Law: An Open Source Chinese Legal Large Language Model
Zhiwei Fei, Songyang Zhang, Xiaoyu Shen +9
While large language models (LLMs) have showcased impressive capabilities, they struggle with addressing legal queries due to the intricate complexities and specialized expertise r…
MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark
Hongwei Liu, Zilong Zheng, Yuxuan Qiao +7
Recent advancements in large language models (LLMs) have showcased significant improvements in mathematics. However, traditional math benchmarks like GSM8k offer a unidimensional p…
InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning
Huaiyuan Ying, Shuo Zhang, Linyang Li +19
The math abilities of large language models can represent their abstract reasoning ability. In this paper, we introduce and open-source our math reasoning LLMs InternLM-Math which…
LawBench: Benchmarking Legal Knowledge of Large Language Models
Zhiwei Fei, Xiaoyu Shen, Dawei Zhu +6
Large language models (LLMs) have demonstrated strong capabilities in various aspects. However, when applying them to the highly specialized, safe-critical legal domain, it is uncl…