Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Heuristics for spotting AI writing, like the overuse of em dashes, are becoming obsolete as models learn from human feedback. For instance, ChatGPT now uses em dashes *less* frequently than human writers at The Economist, flipping the old tell on its head and complicating detection efforts.

Related Insights

OpenAI has publicly acknowledged that the em-dash has become a "neon sign" for AI-generated text. They are updating their model to use it more sparingly, highlighting the subtle cues that distinguish human from machine writing and the ongoing effort to make AI outputs more natural and less detectable.

Medium's platform automatically converted double hyphens to em dashes for years, a stylistic preference of founder Evan Williams. This saturated its content with the punctuation mark, causing AI models trained on its vast corpus to replicate this quirk, effectively becoming a "tell" for AI-generated text.

Current AI models often provide long-winded, overly nuanced answers, a stark contrast to the confident brevity of human experts. This stylistic difference, not factual accuracy, is now the easiest way to distinguish AI from a human in conversation, suggesting a new dimension to the Turing test focused on communication style.

Pangram Labs' detector isn't hard-coded. It's a deep learning model trained on millions of examples. For each human text (e.g., a Yelp review), it sees an AI-generated equivalent, learning the subtle, often inarticulable, differences in word choice and structure that separate them.

Historically, well-structured writing served as a reliable signal that the author had invested time in research and deep thinking. Economist Bernd Hobart notes that because AI can generate coherent text without underlying comprehension, this signal is lost. This forces us to find new, more reliable ways to assess a person's actual knowledge and wisdom.

The tendency for AI models to overuse em dashes may stem from their training data. To expand their knowledge, companies digitized millions of older books, including 19th-century classics where dash usage was at its historical peak. The models simply adopted this stylistic habit.

Once a staple of human literary expression, the em dash is now often perceived as a sign of AI-generated content. This shift has led to writers, like journalist Brian Vance, being wrongly accused of using AI, highlighting a new form of digital misinterpretation.

AI-generated text often uses devices like em-dashes or structuring ideas in threes. These aren't random; they're patterns learned from scraping skilled human writers like C.S. Lewis. This creates a paradox where the stylistic habits of good writing can now be misinterpreted as tells for AI.

While the em dash is a known sign of AI writing, a more subtle indicator is "contrastive parallelism"—the "it's not this, it's that" structure. This pattern, likely learned from marketing copy, is frequently used by LLMs but is uncommon in typical human writing.

When a brand like Apple has a massive, stylistically consistent public corpus, LLMs become experts at mimicking it. This creates a paradox where new, human-written content is flagged as AI-generated because detectors recognize the perfectly emulated patterns they were trained on.