3 papers
cs.CL2025
Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim Evidence Reasoning
Shashidhar Reddy Javaji, Yupeng Cao, Haohang Li +3
Large language models (LLMs) are increasingly being used for complex research tasks such as literature review, idea generation, and scientific paper analysis, yet their ability to…
cs.CE2024
INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent
Haohang Li, Yupeng Cao, Yangyang Yu +12
Recent advancements have underscored the potential of large language model (LLM)-based agents in financial decision-making. Despite this progress, the field currently encounters tw…
cs.DC2024
Kilometer-Level Coupled Modeling Using 40 Million Cores: An Eight-Year Journey of Model Development
Xiaohui Duan, Yuxuan Li, Zhao Liu +38
With current and future leading systems adopting heterogeneous architectures, adapting existing models for heterogeneous supercomputers is of urgent need for improving model resolu…