César de la Fuente's lab at the University of Pennsylvania is mining living and extinct genomes using OpenAI's Codex and ChatGPT to find antimicrobial peptides capable of killing drug-resistant bacteria. The target is a genuine crisis: antimicrobial resistance kills an estimated 1.27 million people annually, and the antibiotic pipeline has been largely dry for decades.

The method is what makes this worth reading in full. De la Fuente's team treats genome sequences as a kind of biological dark matter, using language models to identify candidate molecules that evolution has already tested but humans have never isolated or synthesized. They have pulled candidates from extinct organisms, including the woolly mammoth, a move that reframes ancient DNA as a pharmacological library rather than a paleontological curiosity.

The immediate question this raises is reproducibility and scale: how many candidates survive wet-lab validation, and at what hit rate compared to traditional screening. That answer is in the source, and it changes how you think about what large language models are actually useful for in drug discovery.

[READ ORIGINAL →]