Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

AI models average their training data, resulting in generic content. This problem is compounded as new AIs train on internet data that is increasingly populated by previous AI generations' bland output, creating a self-reinforcing feedback loop of mediocrity.

Related Insights

Contrary to the hype, AI isn't a substitute for human thought. It's a powerful pattern-matching tool that consumes vast data. A growing problem is that AI is increasingly training on its own regurgitated output, creating a closed loop that lacks genuine novelty or external grounding.

The internet's value stems from an economy of unique human creations. AI-generated content, or "slop," replaces this with low-quality, soulless output, breaking the internet's economic engine. This trend now appears in VC pitches, with founders presenting AI-generated ideas they don't truly understand.

AI makes it easy to generate mediocre content, shrinking the gap between bad and passable. However, the effort required to create truly good, differentiating content has increased, widening the gap between what is passable and what is excellent, making true differentiation more difficult.

Newer LLMs exhibit a more homogenized writing style than earlier versions like GPT-3. This is due to "style burn-in," where training on outputs from previous generations reinforces a specific, often less creative, tone. The model’s style becomes path-dependent, losing the raw variety of its original training data.

The debate over distilling from other AI models is becoming moot. The internet is now so saturated with AI-generated content ('AI slop') that any new model trained on web data is already, by default, being trained on the outputs of its predecessors. Pure 'human data' is a dwindling resource.

The proliferation of low-quality, AI-generated content is a structural issue that cannot be solved with better filtering. The ability to generate massive volumes of content with bots will always overwhelm any curation effort, leading to a permanently polluted information ecosystem.

LLMs conform to the average of their training data. When used for creative tasks like writing, they act as "memetic conformity machines," sanding off originality and producing work that sounds like everything else—the literal definition of mediocre.

The greatest danger of AI content isn't job loss or bad SEO, but a societal one. Since we consume more brand content than educational material, an internet flooded with AI's 'predictive text' based on what's common could relegate collective human knowledge and creativity to a permanent base level.

AI models are trained on vast datasets of existing knowledge. Like a librarian who has read every book, their answers represent an average of what they have 'read.' This makes AI an aggregator of existing ideas, not a generator of truly novel, outlier concepts.

LLMs function by predicting the most probable next word, effectively averaging out language. Over-relying on them for content creation will systematically strip away the unique aspects of your brand's voice, leading to homogenization and risking a 'dead internet' effect.