Showing cs.CYShow all
2 papers · 1 filter
cs.CY2026
Muse Spark Safety & Preparedness Report
Cristina Menghini, Peter Ney, Hamza Kwisaba +117
Muse Spark is the latest large language model developed by Meta. In this report, we first present evaluations for catastrophic risk domains under Meta's Advanced AI Scaling Framewo…
cs.CY2025
Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark
Jasper Götting, Pedro Medeiros, Jon G Sanders +6
We present the Virology Capabilities Test (VCT), a large language model (LLM) benchmark that measures the capability to troubleshoot complex virology laboratory protocols. Construc…