5 papers
RWGBench: Evaluating Scholarly Positioning in Related Work Generation
Anzhe Xie, Weihang Su, Jiaxin Mao +4
Large language models have shown strong fluency in scientific writing, yet the evaluation of related work generation (RWG) remains limited. Existing RWG evaluations largely inherit…
SurGE: A Benchmark and Evaluation Framework for Scientific Survey Generation
Weihang Su, Anzhe Xie, Qingyao Ai +5
The rapid growth of academic literature makes the manual creation of scientific surveys increasingly infeasible. While large language models show promise for automating this proces…
Evaluating Intelligence via Trial and Error
Jingtao Zhan, Jiahao Zhao, Jiayu Li +7
Intelligence is a crucial trait for species to find solutions within a limited number of trial-and-error attempts. Building on this idea, we introduce Survival Game as a framework…
Scaling Laws For Dense Retrieval
Yan Fang, Jingtao Zhan, Qingyao Ai +4
Scaling up neural models has yielded significant advancements in a wide array of tasks, particularly in language generation. Previous studies have found that the performance of neu…
Prompt Refinement with Image Pivot for Text-to-Image Generation
Jingtao Zhan, Qingyao Ai, Yiqun Liu +5
For text-to-image generation, automatically refining user-provided natural language prompts into the keyword-enriched prompts favored by systems is essential for the user experienc…