8 papers
VeriGeo: Controllable Geometry Question Generation with Numerical and Analytical Verification
Xiaoxian Duan, Zequn Liu, Yingce Xia
Geometry problem generation is useful for AI-assisted education and multimodal mathematical reasoning, but reliable synthesis remains difficult because the problem statement, diagr…
ForeSci: Evaluating LLM Agents for Forward-Looking AI Research Judgment
Qiuyu Tian, Haojie Yin, Yingce Xia +2
AI research often requires decisions before future evidence exists: which bottleneck to attack, which direction to pursue, or where a project should be positioned. We introduce For…
SciCustom: A Framework for Custom Evaluation of Scientific Capabilities in Large Language Models
Yiyang Gu, Junwei Yang, Junyu Luo +15
Large language models (LLMs) are increasingly applied to scientific research, yet existing evaluations often fail to reflect the fine-grained capabilities required in practice. Mos…
CoPA: Benchmarking Personalized Question Answering with Data-Informed Cognitive Factors
Hang Su, Zequn Liu, Chen Hu +3
While LLMs have demonstrated remarkable potential in Question Answering (QA), evaluating personalization remains a critical bottleneck. Existing paradigms predominantly rely on lex…
Deciphering Scientific Reasoning Steps from Outcome Data for Molecule Optimization
Zequn Liu, Kehan Wu, Shufang Xie +5
Emerging reasoning models hold promise for automating scientific discovery. However, their training is hindered by a critical supervision gap: experimental outcomes are abundant, w…
MolChord: Structure-Sequence Alignment for Protein-Guided Drug Design
Wei Zhang, Zekun Guo, Yingce Xia +4
Structure-based drug design (SBDD), which maps target proteins to candidate molecular ligands, is a fundamental task in drug discovery. Effectively aligning protein structural repr…