Publications (5)
BioSecBench-Refusal: A paired metric for performance and alignment in agentic biosecurity risk assessment
Edwin H. Wintermute, Harmon Bhasin, Christina M. Agapakis +13
As AI agents are incorporated into life science workflows, the capabilities that speed discovery might also enable misuse. We present BioSecBench-Refusal, a benchmark for risk iden…
Automated Neuron Labelling Enables Generative Steering and Interpretability in Protein Language Models
Arjun Banerjee, David Martinez, Camille Dang +1
Protein language models (PLMs) encode rich biological information, yet their internal neuron representations are poorly understood. We introduce the first automated framework for l…
Characterizing price index behavior through fluctuation dynamics
Prasanta K. Panigrahi, Sayantan Ghosh, Arjun Banerjee +2
We study the nature of fluctuations in variety of price indices involving companies listed on the New York Stock Exchange. The fluctuations at multiple scales are extracted through…
scBench-Long: Verifiable Benchmarking of Long-Horizon Single-Cell Biology
Ian Diks, Zhen Yang, Arjun Banerjee +2
Single-cell studies require analysts to convert raw measurements into specific biological claims through multi-step workflows and integration of metadata, assay context, and auxili…
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance
Harmon Bhasin, Kevin Flyangolts, Dianzhuo Wang +9
As pathogen genomic surveillance scales, the bottleneck is shifting from data generation to analysis. We present BioSecBench-Surveillance, a verifiable benchmark of 100 evaluations…