6 papers
PEARL: Auditable Repair for Scientific Reasoning Graph Extraction
Bohan Su, Pengze Li, Yuchen Lu +1
Scientific Reasoning Graph Extraction (SRGE) aims to recover explicit links among observations, evidence, intermediate claims, and paper-level conclusions. LLMs can produce graph-l…
PhysUniBench: A Multi-Modal Physics Reasoning Benchmark at Undergraduate Level
Lintao Wang, Encheng Su, Jiaqi Liu +11
Physics problem-solving is a challenging domain for AI models, requiring integration of conceptual understanding, mathematical reasoning, and interpretation of physical diagrams. E…
Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows
Wanghan Xu, Yuhao Zhou, Yifan Zhou +104
Despite advances in scientific AI, a coherent framework for Scientific General Intelligence (SGI)-the ability to autonomously conceive, investigate, and reason across scientific do…
ARCHE: A Novel Task to Evaluate LLMs on Latent Reasoning Chain Extraction
Pengze Li, Jiaqi Liu, Junchi Yu +5
Large language models (LLMs) are increasingly used in scientific domains. While they can produce reasoning-like content via methods such as chain-of-thought prompting, these output…
Mimicking the Physicist's Eye:A VLM-centric Approach for Physics Formula Discovery
Jiaqi Liu, Songning Lai, Pengze Li +12
Automated discovery of physical laws from observational data in the real world is a grand challenge in AI. Current methods, relying on symbolic regression or LLMs, are limited to u…
Dynamic Knowledge Exchange and Dual-diversity Review: Concisely Unleashing the Potential of a Multi-Agent Research Team
Weilun Yu, Shixiang Tang, Yonggui Huang +7
Scientific progress increasingly relies on effective collaboration among researchers, a dynamic that large language models (LLMs) have only begun to emulate. While recent LLM-based…