3 papers
cs.CL2025
UniCBE: An Uniformity-driven Comparing Based Evaluation Framework with Unified Multi-Objective Optimization
Peiwen Yuan, Shaoxiong Feng, Yiwei Li +7
Human preference plays a significant role in measuring large language models and guiding them to align with human values. Unfortunately, current comparing-based evaluation (CBE) me…
cs.CL2025
LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient
Peiwen Yuan, Shaoxiong Feng, Yiwei Li +7
The rapid advancement of large language models (LLMs) has led to a surge in both model supply and application demands. To facilitate effective matching between them, reliable, gene…
cs.CL2024
Instruction Embedding: Latent Representations of Instructions Towards Task Identification
Yiwei Li, Jiayi Shi, Shaoxiong Feng +6
Instruction data is crucial for improving the capability of Large Language Models (LLMs) to align with human-level performance. Recent research LIMA demonstrates that alignment is…