most citedTorchCP: A Python Library for Conformal Prediction

1 citations · 1 across the 14 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

StatABench: Dataset and Framework for Evaluating Statistical Analysis Capabilities of LLMs

Youxin Zhu, Yixuan Ding, Peng Lai +3

Statistical analysis is a broad, complex field requiring both domain knowledge and tool proficiency. While prior work has evaluated large language models (LLMs) in this domain, exi…

cs.CL2026

Towards Fair and Comprehensive Evaluation of Routers in Collaborative LLM Systems

Wanxing Wu, He Zhu, Yixia Li +7

Large language models (LLMs) have achieved success, but cost and privacy constraints necessitate deploying smaller models locally while offloading complex queries to cloud-based mo…

cs.CL2025

Exploring Imbalanced Annotations for Effective In-Context Learning

Hongfu Gao, Feipeng Zhang, Hao Zeng +3

Large language models (LLMs) have shown impressive performance on downstream tasks through in-context learning (ICL), which heavily relies on the demonstrations selected from annot…

cs.CL2025

ChineseSafe: A Chinese Benchmark for Evaluating Safety in Large Language Models

Hengxiang Zhang, Hongfu Gao, Qiang Hu +7

With the rapid development of Large language models (LLMs), understanding the capabilities of LLMs in identifying unsafe content has become increasingly important. While previous w…

cs.CL2025

Fine-tuning can Help Detect Pretraining Data from Large Language Models

Hengxiang Zhang, Songxin Zhang, Bingyi Jing +1

In the era of large language models (LLMs), detecting pretraining data has been increasingly important due to concerns about fair evaluation and ethical risks. Current methods diff…