Contrary to the belief that AI will replace physical labs, Chai Discovery's co-founder argues it will increase demand. By generating higher-quality molecular candidates, AI boosts the return on investment for each lab test, justifying more testing, not less. This mirrors how software productivity tools increased the demand for engineers.
Diffusion models were a breakthrough for protein generation because they reframe the problem. Instead of a one-shot generation, they learn to make many small, iterative refinements ("make it slightly better"). This "time to think" approach proved more effective for complex biological structures than previous methods like VAEs.
Chai Discovery's core philosophy is a direct application of "The Bitter Lesson" to biotech. They prioritize scaling compute, data, and simple models over creating complex, bespoke biological modules, betting that general-purpose learning methods will outperform human-engineered ones at scale.
Chai Discovery built its founding team with AI researchers first. They deliberately waited to hire domain experts like antibody engineers until the AI model had reached a milestone where it could actually tackle antibody design problems. This ensures specialists have an immediate impact and aren't hired ahead of the technology's capabilities.
Chai Discovery simplifies the complexity of biology by abstracting different molecular challenges, like designing antibodies vs. mini-proteins, into mere "prompts" for a unified model. This is analogous to how a large language model can handle both math problems and English homework, enabling broader generalization from a single architecture.
By choosing a partnership model over developing its own drugs, Chai Discovery subjects its AI to a higher bar. Its models must generalize across diverse targets for multiple partners like Pfizer and Eli Lilly, preventing them from creating bespoke solutions for a single problem. This business model forces technical rigor and scalability.
Training a language model to predict the next amino acid in a sequence forces it to learn the protein's 3D structure. To make accurate predictions, the model must understand an amino acid's physical microenvironment, effectively deriving 3D spatial relationships from 1D sequence data alone. This demonstrates emergent capabilities of LLMs in biology.
The paradigm for drug development is shifting from being "first" or "best" in a category to being "last in class." Using AI, the goal is to design a molecule with such intentionality and specificity that it becomes the final, definitive therapeutic for a disease, rendering subsequent improvements unnecessary.
Chai Discovery found that its first model, with 23 submodules, was too complex to iterate on and scale effectively. A core guiding principle became radical simplification, which makes it easier to understand model dynamics and identify promising scaling directions, even in a complex domain like biology.
