Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Large Language Models often produce clunky metaphors because their training data is swamped by vast quantities of low-quality text, like anime fan fiction. The sheer volume of amateur writing can overpower the influence of well-crafted literature, leading to subpar output.

Related Insights

The auto-regressive, next-token-prediction nature of current LLMs is a 'really, really weird way to produce stuff.' True human creativity and writing insight involve knowing precisely when to make an unpredictable, non-obvious move. This is directly contrary to the model's core process, which is a slave to its immediate context and favors predictable outputs.

AI models fail at great literary writing because they lack an authentic "voice." This voice isn't just a stylistic quirk; it's the product of an individual's unique life experiences and perspective. Since AI lacks this grounding, its writing feels inauthentic, like an imitation of a style without the substance behind it.

AI makes it easy to generate mediocre content, shrinking the gap between bad and passable. However, the effort required to create truly good, differentiating content has increased, widening the gap between what is passable and what is excellent, making true differentiation more difficult.

Newer LLMs exhibit a more homogenized writing style than earlier versions like GPT-3. This is due to "style burn-in," where training on outputs from previous generations reinforces a specific, often less creative, tone. The model’s style becomes path-dependent, losing the raw variety of its original training data.

Sam Altman acknowledged that models are becoming "spiky," with capabilities improving unevenly. OpenAI intentionally prioritized making GPT-5.2 excel at reasoning and coding, which led to a degradation in its creative writing and prose. This highlights the trade-offs inherent in current model training.

AI excels at replicating patterns from its training data. However, top-tier authors provide value by subverting expectations and introducing surprising connections—a skill rooted in creative, pattern-breaking thought that AI struggles with. The act of writing is the act of thinking, which can't be outsourced.

Luis von Ahn highlights a critical flaw in AI: it generates impressive one-off examples but struggles with quality consistency at production scale. Generating 1,000 stories, for example, reveals a high percentage of "pure slump," requiring intense human oversight to maintain brand quality.

AI models produce poor creative writing because they are trained to optimize for superficial proxies for quality, like the number of metaphors. This 'reward hacking' caters to quick judgments from human evaluators on leaderboards, mistaking flashy complexity for genuine literary taste.

AI-generated text often uses devices like em-dashes or structuring ideas in threes. These aren't random; they're patterns learned from scraping skilled human writers like C.S. Lewis. This creates a paradox where the stylistic habits of good writing can now be misinterpreted as tells for AI.

AI models are trained on vast datasets of existing knowledge. Like a librarian who has read every book, their answers represent an average of what they have 'read.' This makes AI an aggregator of existing ideas, not a generator of truly novel, outlier concepts.