4 papers
Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery
Tianyun Zhong, Wangyi Jiang, Wei Wang +15
Large language models (LLMs) excel at answering pre-specified questions, yet their ability to navigate the open-ended, pre-conclusion stage of discovery remains largely unmeasured.…
DLawBench: Evaluating LLMs Through Multi-Turn Legal Consultation
Li Zhang, Yuzhen Shi, Yiran Hu +15
Lawyer-client consultation is a critical starting point for legal services. Effective legal assistance hinges on eliciting sufficient and truthful information from clients in order…
Socratic-Zero : Bootstrapping Reasoning via Data-Free Agent Co-evolution
Shaobo Wang, Zhengbo Jiao, Zifan Zhang +6
Recent breakthroughs in large language models (LLMs) on reasoning tasks rely heavily on massive, high-quality datasets-typically human-annotated and thus difficult to scale. While…
SKYLENAGE Technical Report: Mathematical Reasoning and Contest-Innovation Benchmarks for Multi-Level Math Evaluation
Hu Wei, Ze Xu, Boyu Yang +15
Large language models (LLMs) now perform strongly on many public math suites, yet frontier separation within mathematics increasingly suffers from ceiling effects. We present two c…