4 papers
MatrAIx: Simulating the World with 8.3 Billion Persona Agents
Xiaomin Li, Yuexing Hao, Jianheng Hou +90
Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and inter…
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance
Harmon Bhasin, Kevin Flyangolts, Dianzhuo Wang +9
As pathogen genomic surveillance scales, the bottleneck is shifting from data generation to analysis. We present BioSecBench-Surveillance, a verifiable benchmark of 100 evaluations…
BioSecBench-Refusal: A paired metric for performance and alignment in agentic biosecurity risk assessment
Edwin H. Wintermute, Harmon Bhasin, Christina M. Agapakis +13
As AI agents are incorporated into life science workflows, the capabilities that speed discovery might also enable misuse. We present BioSecBench-Refusal, a benchmark for risk iden…
Viral Proteins Reveal Geometry of Protein Language Models
Arthur Bigot, Harmon Bhasin, Core Francisco Park +2
Protein language models are trained on highly imbalanced datasets, raising the question of how they represent underrepresented biological sequences. Using viral proteins as a case…