I published a preprint last week. It mostly focuses on what biosecurity evaluations of AI models measure and how far those measurements sit from what you might call an estimate of real-world risk.
AI Biosecurity Benchmarks and Real-World Risk (2026). Available at SSRN: https://ssrn.com/abstract=7358438 or dx.doi.org/10.2139/ssrn.7358438
I’ve spent about 15 years in and adjacent to biodefense. At first as an academic researcher when I got my first faculty position here in the School of Medicine, then for many years at Signature Science where the badge said industry but the work was for the US government. Now I’m back in academia at the UVA School of Data Science, building out our own AIxBio and biosecurity research program along with other colleagues here at SDS and in our National Security Data and Policy Institute.
Over the past year since I returned to UVA I’ve been closely following the AIxBio and biosecurity literature: new benchmarks/evals, human uplift trials, synthesis screening results, unlearning and pretraining-filter experiments, developer preparedness frameworks, policy proposals. My thinking was that the many pages of notes on the important work happening here could be a starting point for developing a really interesting upper level elective or graduate course here in SDS. Many of these notes became posts here on this newsletter. The rest piled up in my Zotero library, full of notes and highlights attached to each of the entries. I decided to package up everything I had into a single review, mostly on the theme of how benchmark scores connect to real-world risk and what work remains to be done. I started writing this mostly as a primer for new people joining my lab, but then figured it may be of broader interest.
The list of groups doing this work or funding it is long, and the reference list includes work from Active Site, SecureBio, Sentinel Bio, Signature Science, IBBIS, LatchBio, RAND Center on AI, Security, and Technology (CAST), Hewlett Foundation, Johns Hopkins Center for Health Security, METR, Epoch AI, UK AI Security Institute, Frontier Model Forum, Nuclear Threat Initiative, International Gene Synthesis Consortium, Sequence Biosecurity Risk Consortium, Aclid, Centre for Long-Term Resilience, NIST Center for AI Standards and Innovation (CAISI), Georgetown CSET, all the frontier labs, and many, many others.
I don’t love offering a TL;DR on a review paper since synthesizing all the interesting bits from hundreds of papers is kind of the whole point of writing one. But it might just be TL, so in case you DR, here’s a one-line summary:
Really great people at really great organizations are doing really great work at the intersection of AI and biosecurity.
I tried to organize as much of it as I could into one place, with sincere apologies to anyone whose work in this space I omitted.1 If you’d be willing to peer review this when I try to publish it, email me.
And even more sincere apologies to any work I cited here but mischaracterized. The last page of the paper has the most detailed AI disclosure I’ve ever read or written. On an early draft I gave an army of Claude agents the manuscript Quarto file, full access to my Zotero library and PDF exports alongside the BiBTeX file with everything I cited, and asked for rigorous fact checking of everything I wrote and every statistic I cited. It worked well, until it didn’t — because of the nature of the subject matter being biosecurity and discussion of dual-use technology, adversarial red-teaming, safeguards, screening evasion, etc., I ended up hitting refusals from nearly every model I attempted to use in this fact-checking agentic workflow. I eventually gave up and did the fact-checking and proofreading the old fashioned way (re-reading, and re-re-reading), but I’m human and I make human mistakes. Hopefully anything I didn’t catch will get surfaced during peer review. See the last page of the paper for details.

