4 papers
Decoupled Mixture-of-Experts for Parametric Knowledge Injection
Baoqing Yue, Weihang Su, Qingyao Ai +5
Knowledge injection aims to equip large language models (LLMs) with external, domain-specific, or time-sensitive knowledge. Existing approaches typically face a trade-off between f…
Interactive Benchmarks
Baoqing Yue, Zihan Zhu, Yutong Han +6
Existing reasoning evaluation paradigms suffer from different limitations: fixed benchmarks are increasingly saturated and vulnerable to contamination, while preference-based evalu…
Relative-Based Scaling Law for Neural Language Models
Baoqing Yue, Jinyuan Zhou, Zixi Wei +3
Scaling laws aim to accurately predict model performance across different scales. Existing scaling-law studies almost exclusively rely on cross-entropy as the evaluation metric. Ho…
JuDGE: Benchmarking Judgment Document Generation for Chinese Legal System
Weihang Su, Baoqing Yue, Qingyao Ai +6
This paper introduces JuDGE (Judgment Document Generation Evaluation), a novel benchmark for evaluating the performance of judgment document generation in the Chinese legal system.…