2 papers
cs.AI2026
K-Bench: measuring model performance on real scientific agent requests
Aubrey Brueckner, Darshil Patel, Yuhuan He +1
Benchmarks for scientific artificial intelligence are mostly written to be scored: multiple-choice questions, curated agent tasks with reference solutions, or simulators with a kno…
cs.AI2025
K-Dense Analyst: Towards Fully Automated Scientific Analysis
Orion Li, Vinayak Agarwal, Summer Zhou +2
The complexity of modern bioinformatics analysis has created a critical gap between data generation and developing scientific insights. While large language models (LLMs) have show…