Early AI text had recognizable patterns or "tells," which machine learning classifiers could detect. As generative models improve, these tells disappear, causing the accuracy of such classifiers to plummet over time. This makes active watermarking a more robust solution than passive detection.
Watermarking isn't about hiding data in whitespace. It works by using a secret key to subtly bias an LLM's selection of the next word from a list of valid options. A detector can then identify this biased pattern across a body of text to confirm its AI origin without altering meaning.
Watermarking isn't a one-size-fits-all solution. High-dimensional data like images (millions of pixels) offers ample space to hide an imperceptible signal. In contrast, low-dimensional data like text requires a completely different method based on biasing word choice to avoid altering its meaning.
To validate their watermarking method, researchers didn't just rely on simulations. They physically created watermarked protein binders in a lab. Lab tests showed that watermarked proteins had nearly identical "hit rates" and binding capabilities as non-watermarked versions, proving the method's real-world effectiveness.
While making detection tools public seems beneficial, it introduces a significant risk. Adversaries can repeatedly query the detector with different inputs to learn its patterns and eventually reverse-engineer the secret key. This would allow them to either remove watermarks or fool that specific detector.
A key biosecurity risk is that AI can generate a protein sequence that looks innocent and doesn't match known threats in a screening database. However, this novel sequence can fold into a 3D structure with the same harmful function as a restricted pathogen, effectively sneaking past current safety checks.
DeepMind’s SynthID Bio embeds watermarks in protein sequences by substituting certain amino acids with functionally similar ones (e.g., leucine for isoleucine). This is analogous to choosing a synonym in text. The change is imperceptible to the protein's function but detectable as a watermarked pattern.
The goal for robust watermarking is to embed the signal so deeply that any attempt to remove it also fundamentally degrades the quality or purpose of the original content. For an image, this means destroying its visual integrity; for a protein, it means altering its biological function.
