1 citations · 1 across the 14 of their papers we have counts for
5 papers · 1 filter
StatABench: Dataset and Framework for Evaluating Statistical Analysis Capabilities of LLMs
Youxin Zhu, Yixuan Ding, Peng Lai +3
Statistical analysis is a broad, complex field requiring both domain knowledge and tool proficiency. While prior work has evaluated large language models (LLMs) in this domain, exi…
Towards Fair and Comprehensive Evaluation of Routers in Collaborative LLM Systems
Wanxing Wu, He Zhu, Yixia Li +7
Large language models (LLMs) have achieved success, but cost and privacy constraints necessitate deploying smaller models locally while offloading complex queries to cloud-based mo…
Exploring Imbalanced Annotations for Effective In-Context Learning
Hongfu Gao, Feipeng Zhang, Hao Zeng +3
Large language models (LLMs) have shown impressive performance on downstream tasks through in-context learning (ICL), which heavily relies on the demonstrations selected from annot…
ChineseSafe: A Chinese Benchmark for Evaluating Safety in Large Language Models
Hengxiang Zhang, Hongfu Gao, Qiang Hu +7
With the rapid development of Large language models (LLMs), understanding the capabilities of LLMs in identifying unsafe content has become increasingly important. While previous w…
Fine-tuning can Help Detect Pretraining Data from Large Language Models
Hengxiang Zhang, Songxin Zhang, Bingyi Jing +1
In the era of large language models (LLMs), detecting pretraining data has been increasingly important due to concerns about fair evaluation and ethical risks. Current methods diff…