2 citations · 2 across the 12 of their papers we have counts for
7 papers · 1 filter
PEARL: Auditable Repair for Scientific Reasoning Graph Extraction
Bohan Su, Pengze Li, Yuchen Lu +1
Scientific Reasoning Graph Extraction (SRGE) aims to recover explicit links among observations, evidence, intermediate claims, and paper-level conclusions. LLMs can produce graph-l…
LECTOR: Joint Optimization of Scientific Reasoning Graphs and Introduction Generation
Jiabei Xiao, Yizhou Wang, Chen Tang +3
AI Scientists have shown promising progress across multiple stages of the research pipeline, among which automatic scientific paper writing remains a formidable challenge. The Intr…
AI-for-Science Low-code Platform with Bayesian Adversarial Multi-Agent Framework
Zihang Zeng, Jiaquan Zhang, Pengze Li +2
Large Language Models (LLMs) demonstrate potentials for automating scientific code generation but face challenges in reliability, error propagation in multi-agent workflows, and ev…
PhysUniBench: A Multi-Modal Physics Reasoning Benchmark at Undergraduate Level
Lintao Wang, Encheng Su, Jiaqi Liu +11
Physics problem-solving is a challenging domain for AI models, requiring integration of conceptual understanding, mathematical reasoning, and interpretation of physical diagrams. E…
SciIF: Benchmarking Scientific Instruction Following Towards Rigorous Scientific Intelligence
Encheng Su, Jianyu Wu, Chen Tang +9
As large language models (LLMs) transition from general knowledge retrieval to complex scientific discovery, their evaluation standards must also incorporate the rigorous norms of…
ARCHE: A Novel Task to Evaluate LLMs on Latent Reasoning Chain Extraction
Pengze Li, Jiaqi Liu, Junchi Yu +5
Large language models (LLMs) are increasingly used in scientific domains. While they can produce reasoning-like content via methods such as chain-of-thought prompting, these output…