Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents
Jinge Wu, Hongjian Zhou, Mingde Zeng +8
Reproducing and comparing deep research agents today is hard: the same backbone evaluated on the same benchmark can report different accuracies across papers because the harness an…
cs.AI2026
Scientific reasoning does not reliably translate into scientific forecasting in frontier AI
Sean Wu, Pan Lu, Yupeng Chen +7
AI systems are increasingly used to support forward-looking scientific judgment, but it remains unclear whether they can form reliable expectations about future scientific advances…