DeepMind Watermarks AI-Designed Proteins for Biosecurity

Google DeepMind introduced SynthID Bio, a system that embeds detectable watermarks into AI-generated protein sequences and structures while preserving…

Ayla Demirhan ·

DeepMind Watermarks AI-Designed Proteins for Biosecurity

Google DeepMind researchers have developed a method to embed detectable watermarks into artificial intelligence-generated protein sequences and structures, preserving their biological function. This initiative, named SynthID Bio, extends provenance systems from AI-generated media into synthetic biology, addressing concerns about the origin of novel biological designs.

Watermarking AI-Generated Biological Sequences

The SynthID Bio system integrates a unique signature directly into a protein's amino acid sequence or its predicted three-dimensional structure. Unlike image watermarking, where minor pixel changes are often imperceptible, altering a protein can fundamentally change its folding, binding capabilities, or overall function. DeepMind's research focused on ensuring watermarked biological designs remained functional.

Functional Integrity Confirmed in Lab Tests

Laboratory tests, detailed in a Nature report, demonstrated that watermarked protein binders maintained functionality across three specific targets: VEGF-A, PD-L1, and the SARS-CoV-2 spike protein's receptor-binding domain. Pushmeet Kohli, Chief Scientist at Google Cloud and VP Science at Google DeepMind, affirmed the successful synthesis of functional, watermarked AI-designed proteins.

Dual Embedding Approach for Proteins

SynthID Bio employs two distinct methods. For protein sequences, the system subtly modifies amino acid selection during the AI model's generation, allowing detection with a secret key. For three-dimensional protein structures, researchers fine-tuned a component of AlphaFold 3, embedding the watermark directly into predicted atomic coordinates. The structural method achieved a detection rate exceeding 99.8% with a 0.1% false-positive rate in tests, without compromising structural accuracy compared to the AlphaFold 3 baseline.

Addressing Emerging Biosecurity Challenges

SynthID Bio responds to generative AI's increasing capacity to create novel biological sequences, which complicates existing biosecurity protocols. Current DNA synthesis companies screen orders against databases of known threats. However, a completely new AI-generated sequence might not resemble cataloged threats, making its benign nature difficult to ascertain without further review.

Enhancing Provenance Tracking for Biologics

SynthID Bio aims to provide an additional signal for biosecurity. If a synthesis provider verifies an unfamiliar sequence originated from a known AI system with established safeguards, this information could integrate into the screening process. Sarah Carter, Principal at Science Policy Consulting, characterized this as a crucial element for tracking biological design provenance. Similarly, large volumes of unlabeled AI-generated material entering scientific databases like the Protein Data Bank or GenBank could undermine data integrity, a concern a detectable watermark could mitigate by flagging AI-generated entries.

Limitations and Future Development Outlook

Despite its potential, SynthID Bio is a proof of concept, requiring further technical development, coordination, and standardization for operational use in biosecurity and scientific databases. The research also identified specific weaknesses in current watermarking approaches. The sequence watermark can be removed through resequencing with ProteinMPNN, a process that stripped the watermark in tests involving 38,396 binders, though potentially at a cost to protein function. The structure watermark, while resilient to minor digital noise, could be destroyed by a constrained structural relaxation process.

Zero-Bit Watermark Functionality and Scope

Both current SynthID Bio versions utilize a zero-bit watermark, meaning the detector only confirms a watermark's presence, not richer information like user identification or additional provenance data. This limits its utility for distinguishing multiple users or encoding complex metadata. The Alphabet-funded study involved Alphabet employees as authors, who may hold company stock. DeepMind has made its methods paper, code, and in vitro data publicly available, alongside access information for the recommended SynthID Bio structure model weights, with ongoing collaborations to extend the work beyond individual proteins.

More stories

Latest news