Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
LabOSBench: Benchmarking Computer Use Agents for Scientific Instrument Control
Anqi Zou, Han Deng, Chengyu Zhang +9
Current computer-use benchmarks primarily focus on software operation tasks in virtualized systems, whereas scientific instrumentation scenarios require coordinated control over co…
cs.AI2026
SciEvalKit: An Open-source Evaluation Toolkit for Scientific General Intelligence
Yiheng Wang, Yixin Chen, Shuo Li +33
We introduce SciEvalKit, a unified benchmarking toolkit designed to evaluate AI models for science across a broad range of scientific disciplines and task capabilities. Unlike gene…