Screening for Function, not just Sequence
A proposed successor to IARPA's FunGCAT program
Synthetic DNA providers typically screen orders by matching sequences against databases of known pathogens. Why? A “dangerous” sequence has, until recently, looked like something already cataloged. Protein design models make this trickier. Wittmann et al. at Microsoft, Battelle, IDT, Twist, and others showed last year that AI could rewrite known toxins into synthetic homologs with the same function and a different amino acid sequence, and that these paraphrased versions slipped past existing screening tools.
Hanna Palya and Cassidy Nelson propose ADAPT (Advanced Detection of AI-enabled Pathogenic Threats) as a five-year successor to the IARPA FunGCAT program that ran from 2017 to 2022.
Palya H and Nelson C (2026) ADAPT: a programme for the advanced detection of AI-enabled pathogenic threats. Front. Bioeng. Biotechnol. 14:1819372. doi: 10.3389/fbioe.2026.1819372
The case for ADEPT is that the field has no operationalizable definition of what makes a sequence concerning,1 which leaves providers interpreting vague guidance and regulators unable to enforce some floor for “concern.” And when current tools detect threats by sequence similarity rather than function, a redesigned virulence factor might slip past these filters, being read as benign.
As proposed, Phase I would replace static watchlists with a weighted, hierarchical ontology that scores sequences across attributes like molecular function, necessity for pathogenicity, host range, and weaponization potential. A housekeeping gene shared with a pathogen scores low, but a virulence factor essential for host harm would score high, even after its sequence has been redesigned past recognition. Phase II would build the tools as cascades, where cheap filters screen everything and expensive protein language model analysis handles only what makes it through to this point.

The hard part (as usual) is annotation. Pathogenic functions might sit on peripheral nodes of a Gene Ontology graph, and CAFA benchmarks show function predictors aren’t that great there. Palya and Nelson lean on protein language model embeddings, few-shot classifiers, and graph neural networks over GO to populate the ontology faster than manual curation allows for. They cite few-shot work that generalized to new toxin classes from tens of examples, which would let screening begin before large training sets exist. Budget estimate: $65 to $80 million, i.e., between an ARIA and a DARPA program.
Full disclosure: I worked on the IARPA FunGCAT program for a few years with SigSci and friends. Throat-clearing out of the way: if you’re listening out there, IARPA (and I’m sure you are), I’d love to see a FunGCAT follow-on program like the ADEPT proposal here get some attention (read: $$$) from the likes of IARPA, DARPA, or even a potential DOE Genesis Mission-like program.
Although maybe we are moving toward an operationalizable definition of a sequence of concern. See the June 3 2026 paper from some of the usual suspects: Alexanian T, Beal J, Bartling C, Berlips J, Carr PA, Clore A, Cozzarini H, Diggans J, El Moubayed Y, Esvelt K, Flyangolts K, Foner L, Fullerton PA, Gemler BT, Jagla CAD, Lababidi R, Mitchell T, Murphy ST, Parker MT, Roehner N, Rusch A, Talley K, Timmerman T and Wheeler NE (2026) Developing a standard definition for sequences of concern. Front. Bioeng. Biotechnol. 14:1830842. doi: 10.3389/fbioe.2026.1830842.
