1 citations · 1 across the 2 of their papers we have counts for
4 papers
RWGBench: Evaluating Scholarly Positioning in Related Work Generation
Anzhe Xie, Weihang Su, Jiaxin Mao +4
Large language models have shown strong fluency in scientific writing, yet the evaluation of related work generation (RWG) remains limited. Existing RWG evaluations largely inherit…
MetaSyn: A Benchmark for LLM Agents on Meta-Analysis Articles from Nature Portfolio
Anzhe Xie, Weihang Su, Yujia Zhou +3
Systematic review and meta-analysis is an important method for scientific research. It comprehensively studies target research questions by combining evidence from multiple indepen…
SurGE: A Benchmark and Evaluation Framework for Scientific Survey Generation
Weihang Su, Anzhe Xie, Qingyao Ai +5
The rapid growth of academic literature makes the manual creation of scientific surveys increasingly infeasible. While large language models show promise for automating this proces…
STARD: A Chinese Statute Retrieval Dataset with Real Queries Issued by Non-professionals
Weihang Su, Yiran Hu, Anzhe Xie +6
Statute retrieval aims to find relevant statutory articles for specific queries. This process is the basis of a wide range of legal applications such as legal advice, automated jud…