Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Contrary to the norm, TypeSafe AI avoids training on user data. They believe real-world data is heavily biased towards current use cases, which would cause the model to "fracture" and fail on future, unimagined applications. Their goal is a general cognitive core, not a model optimized for today's queries.

Related Insights

Resolve AI trains its production debugging models not on private customer code, but on the sequence of actions humans take to solve problems. This involves long, multi-step tasks across systems like Datadog and AWS—a type of data that general-purpose models lack.

Despite processing 15 million clinical charts, Datycs doesn't use this data for model training. Their agreements explicitly respect that data belongs to the patient and the client—an ethical choice that prevents them from building large, aggregated language models from customer data.

A key pillar of human-centric AI is ensuring data is "future-proof." Because models are trained on historical data, they can quickly become irrelevant or harmful as market conditions change. This requires a proactive strategy to prevent model decay, not just reactive fixes after failures occur.

Microsoft's case management AI avoids training directly on private customer data. Instead, it operates on a "bring your own knowledge" model, using only the knowledge articles and resources explicitly provided by the customer. This approach sidesteps major privacy and data governance concerns common in enterprise AI adoption.

To lower the activation energy for user adoption, OpenAI deliberately will not use data connected to ChatGPT Health to train its foundation models. This strategic choice is designed to remove any tension between privacy and utility, assuring users their sensitive information is not being used for other purposes and building the trust necessary for scaled impact in the healthcare domain.

Diogo Almeida claims that even with a billion-dollar investment, he would not engage in pre-training a new foundation model. He believes the most significant leverage and innovation comes from post-training techniques and superior data strategy, which can create more value than competing on raw compute for pre-training.

Synthetic models don't merely inherit human biases because they are trained on vast datasets that have already been processed, scrubbed, and validated by researchers. The AI learns from the 'corrected' view of public opinion, not the raw, biased inputs from individual survey takers.

Instead of costly proprietary data generation, Turbine focused on the 'unsexy' work of combining many different public and partner datasets. This capital-efficient approach forced them to build an AI model architected for generalization and data efficiency from the very beginning.

Contrary to the goal of perfect data retention, 'machine unlearning' is becoming a critical capability. The ability for an AI to forget is essential for privacy (removing user data), correcting biases from flawed training data, and adapting to new information, mirroring a core, beneficial aspect of human cognition.

AI products in easily verifiable domains (like coding) will be dominated by large labs. Startup defensibility lies in generating unique data where success is hard to verify automatically, like genuine human learning. This requires real user session data to create a data flywheel that frontier models cannot replicate.

TypeSafe AI Intentionally Avoids User Data to Prevent Overfitting to the Present | RiffOn