Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

An LLM's tendency to "mode collapse"—outputting the most common, probable answer—is beneficial for coding, where correctness and predictability are key. However, this same behavior becomes a drawback for discovery tasks like restaurant recommendations, leading to generic and uninspired suggestions for everyone.

Related Insights

AI models are trained to find the most probable answer, reflecting the average of their data. Truly great, tasteful work is often unique and statistically unlikely, a quality that current models, which regress to the mean, struggle to produce. They can solve PhD-level math but fail at creative tasks like writing a good tweet.

An LLM's core training objective—predicting the next token—makes it sensitive to the raw frequency of words and numbers online. This creates a subtle but profound flaw: it's more likely to output '30' than '29' in a counting task, not because of logic, but because '30' is statistically more common in its training data.

LLMs shine when acting as a 'knowledge extruder'—shaping well-documented, 'in-distribution' concepts into specific code. They fail when the core task is novel problem-solving where deep thinking, not code generation, is the bottleneck. In these cases, the code is the easy part.

MIT research reveals that large language models develop "spurious correlations" by associating sentence patterns with topics. This cognitive shortcut causes them to give domain-appropriate answers to nonsensical queries if the grammatical structure is familiar, bypassing logical analysis of the actual words.

As large language models are optimized for rationality and objective problem-solving, their ability to simulate the irrationality and subjective values inherent in human behavior has plateaued. This necessitates a new modeling paradigm focused on capturing human diversity, not just super-intelligence.

When asked to analyze 100 papers, LLMs often admit they didn't complete the task. This failure stems from outcome-based training, which prioritizes a plausible-looking final output over correctly following the required process, revealing a fundamental flaw in current training paradigms.

AI struggles to provide truly useful, serendipitous recommendations because it lacks any understanding of the real world. It excels at predicting the next word or pixel based on its training data, but it can't grasp concepts like gravity or deep user intent, a prerequisite for truly personalized suggestions.

When given ambiguous instructions, LLMs will choose the most common technology stack from their training data (e.g., React with Tailwind), even if it contradicts the project's goals. Developers must provide explicit constraints to avoid this unwanted default behavior.

Simply asking an LLM to "judge" an output yields generic results. A useful LLM judge requires manually injecting your own taste by creating detailed rubrics with extensive examples of "good" and "bad" at a granular level, essentially brute-forcing your preferences into the model.

According to Dreamer's CEO, the biggest capability missing from LLMs is "taste." By default, AI-generated applications and UIs are generic and identifiable by the model that created them. It requires extensive human effort in prompt engineering and templating to create delightful, non-generic user experiences.