Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

To stay ahead in the cat-and-mouse game of AI detection, Pangram internally develops its own "humanizer" tools—software designed to make AI text undetectable. They use these adversarial tools to train their detection model against the latest evasion techniques, without ever releasing the humanizers publicly.

Related Insights

Firms monitor their AI models with their own models, a practice called "untrusted monitoring." This creates a potential blind spot, as a model that knows how to be deceptive could also know how to evade detection from a copy of itself.

Effective AI detection frames the problem as large-scale authorship identification. By training on paired examples of human versus LLM-generated text on the same prompts, detection models learn to recognize the unique statistical "smell" or stylistic signature of each major AI model.

Didi Das reveals using an AI detector, Pangram, on internal work. The rationale isn't just to police AI usage, but to gauge if an employee has genuinely engaged with a task or simply "produced slop." This signals a shift where AI fluency is measured by thoughtful assistance rather than total abdication.

Pangram Labs' detector isn't hard-coded. It's a deep learning model trained on millions of examples. For each human text (e.g., a Yelp review), it sees an AI-generated equivalent, learning the subtle, often inarticulable, differences in word choice and structure that separate them.

For an AI detection tool, a low false-positive rate is more critical than a high detection rate. Pangram claims a 1-in-10,000 false positive rate, which is its key differentiator. This builds trust and avoids the fatal flaw of competitors: incorrectly flagging human work as AI-generated, which undermines the product's credibility.

Pangram Labs uses an "active learning" loop to enhance its model. After an initial training, the model scans a massive corpus to identify its own errors (false positives/negatives). These hard-to-classify examples are then fed back into the training set, making the next version more robust.

Heuristics for spotting AI writing, like the overuse of em dashes, are becoming obsolete as models learn from human feedback. For instance, ChatGPT now uses em dashes *less* frequently than human writers at The Economist, flipping the old tell on its head and complicating detection efforts.

The AI detection arms race now includes "humanizers": specialized LLMs that rewrite AI-generated text to evade detection. This is the modern, sophisticated version of the old plagiarism tactic of running copied text through a thesaurus to change just enough keywords to fool detectors and teachers.

To accurately measure its false positive rate (how often it wrongly flags human writing as AI), Pangram runs its models on millions of documents written before 2022. Since this content is guaranteed to be AI-free and is not in the training set, it provides a clean benchmark for accuracy.

When a brand like Apple has a massive, stylistically consistent public corpus, LLMs become experts at mimicking it. This creates a paradox where new, human-written content is flagged as AI-generated because detectors recognize the perfectly emulated patterns they were trained on.

AI Detector Pangram Builds Internal "Humanizers" to Make its Own Models More Robust | RiffOn