Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Unlike models like Claude which can be overly verbose, the Grok model's writing is described as "terrible" due to being too brief and token-efficient. It unnaturally compresses sentences into a few words, making it hard to understand and requiring specific voice training.

Related Insights

Beyond raw capability, top AI models exhibit distinct personalities. Ethan Mollick describes Anthropic's Claude as a fussy but strong "intellectual writer," ChatGPT as having friendly "conversational" and powerful "logical" modes, and Google's Gemini as a "neurotic" but smart model that can be self-deprecating.

AI models fail at great literary writing because they lack an authentic "voice." This voice isn't just a stylistic quirk; it's the product of an individual's unique life experiences and perspective. Since AI lacks this grounding, its writing feels inauthentic, like an imitation of a style without the substance behind it.

Newer LLMs exhibit a more homogenized writing style than earlier versions like GPT-3. This is due to "style burn-in," where training on outputs from previous generations reinforces a specific, often less creative, tone. The model’s style becomes path-dependent, losing the raw variety of its original training data.

Sam Altman acknowledged that models are becoming "spiky," with capabilities improving unevenly. OpenAI intentionally prioritized making GPT-5.2 excel at reasoning and coding, which led to a degradation in its creative writing and prose. This highlights the trade-offs inherent in current model training.

Current AI models often provide long-winded, overly nuanced answers, a stark contrast to the confident brevity of human experts. This stylistic difference, not factual accuracy, is now the easiest way to distinguish AI from a human in conversation, suggesting a new dimension to the Turing test focused on communication style.

Sam Altman admitted OpenAI intentionally neglected the model's writing style, which became unwieldy, to focus limited resources on enhancing its core intelligence and engineering capabilities. This reveals a strategy of prioritizing foundational model improvements over user-facing polish during development cycles.

Models like GPT Live prioritize low latency and natural interaction, making them feel more human. However, this is a specific optimization target that differs from deep, strategic reasoning. Users must understand they are interacting with a conversational layer, which may not have the same raw intelligence as the underlying frontier model it calls upon.

When AI labs release new models, they may de-prioritize certain skills like writing to focus on others like agentic capabilities. This causes noticeable shifts in tone and quality, forcing users to re-evaluate and adjust their custom instructions for GPTs and other AI tools.

The speaker coins the term "Claude Slop" for the frustratingly indirect, apologetic, and prose-heavy communication style of Anthropic's models. This specific type of poor output is a major user experience hurdle, making the model difficult to read and act upon, even when the underlying work is high-quality.

Large Language Models often produce clunky metaphors because their training data is swamped by vast quantities of low-quality text, like anime fan fiction. The sheer volume of amateur writing can overpower the influence of well-crafted literature, leading to subpar output.

Grok AI Models Suffer from Overly Clipped, Unnatural Writing, a Contrast to Other Verbose AIs | RiffOn