2 papers
cs.CL2026
Benchmark^2: Systematic Evaluation of LLM Benchmarks
Qi Qian, Chengsong Huang, Jingwen Xu +13
The rapid proliferation of benchmarks for evaluating large language models (LLMs) has created an urgent need for systematic methods to assess benchmark quality itself. We propose B…
cs.AI2025
RECAST: Expanding the Boundaries of LLMs' Complex Instruction Following with Multi-Constraint Data
Zhengkang Guo, Wenhao Liu, Mingchen Xie +13
Large language models (LLMs) are increasingly expected to tackle complex tasks, driven by their expanding applications and users' growing proficiency in crafting sophisticated prom…