AI-generated content, like product reviews, appears high-quality individually but reveals its artificiality in bulk. It defaults to a safe, descriptive "center of the distribution" (mode collapse), lacking the variation and personality of genuine human writing, making it detectable at scale.
Common AI writing clichés, like overusing "delve" or the "it's not just X, but Y" trope, aren't just artifacts of next-word prediction. They are amplified by reinforcement learning (RL), which rewards models for producing text that "sounds" engaging, leading to overuse of certain stylistic devices.
The CEO of Pangram, Max Spero, believes that as AI makes intelligence abundant and nearly free, value will shift to the remaining scarce resource: authentic human input and creativity. His company's mission is to build tools that identify and preserve this value.
To stay ahead in the cat-and-mouse game of AI detection, Pangram internally develops its own "humanizer" tools—software designed to make AI text undetectable. They use these adversarial tools to train their detection model against the latest evasion techniques, without ever releasing the humanizers publicly.
While one could fine-tune a custom AI to evade detection, the most capable models are centralized and expensive. This market concentration means most users rely on a few common models (like from OpenAI or Anthropic), making their distinct "fingerprints" easier for detectors like Pangram to identify.
A subtle hallmark of AI-generated text is its lack of pacing. Unlike human writing, which builds towards a conclusion, AI models often try to make every single sentence sound impressive and important. This creates a tiring and monotonous reading experience that makes a reader's "eyes glaze over."
To accurately measure its false positive rate (how often it wrongly flags human writing as AI), Pangram runs its models on millions of documents written before 2022. Since this content is guaranteed to be AI-free and is not in the training set, it provides a clean benchmark for accuracy.
According to Pangram's CEO, different LLMs have unique "voices." Anthropic's Claude tends to be verbose and hedges statements, while OpenAI's ChatGPT is more curt and favors short, staccato sentences. These stylistic fingerprints allow detection tools to not only identify AI text but also its likely source.
