2 papers
cs.AI2026
SciHazard: A Benchmark for Measuring Scientific Safety Risks with Decomposed Harm Scoring
Chunxiao Li, Yuan Xiong, Lijun Li +4
Large language models (LLMs) increasingly support science, but they can also convert hazardous scientific knowledge into actionable misuse guidance. Existing benchmarks often rely…
cs.CL2025
ATLAS: A High-Difficulty, Multidisciplinary Benchmark for Frontier Scientific Reasoning
Hongwei Liu, Junnan Liu, Shudong Liu +33
The rapid advancement of Large Language Models (LLMs) has led to performance saturation on many established benchmarks, questioning their ability to distinguish frontier models. Co…