1 paper
Shitou Zhang, Zuchao Li, Xingshen Liu +2
In light of the rapidly evolving capabilities of large language models (LLMs), it becomes imperative to develop rigorous domain-specific evaluation benchmarks to accurately assess…