Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Rather than a mysterious aesthetic sense, mathematical "taste" can be defined operationally for an AI. It's the ability to make better judgments that increase the speed and success rate of solving difficult problems. By this utilitarian metric, as AIs solve increasingly harder problems, their "taste" is definitionally improving.

Related Insights

AI models are trained to find the most probable answer, reflecting the average of their data. Truly great, tasteful work is often unique and statistically unlikely, a quality that current models, which regress to the mean, struggle to produce. They can solve PhD-level math but fail at creative tasks like writing a good tweet.

A significant but underappreciated strength of AI in math is its ability to perfectly execute the minute details of an idea. While humans get lost in the 'epsilon smaller than delta' complexities, AI systems consistently and correctly handle these finicky arguments, which is often the primary barrier to proving a result.

The goal for AI should be to surpass human rationality, not merely replicate it. Just as a calculator is designed to be better at arithmetic, AI should be built to overcome our cognitive biases and be superior at manipulating probabilities to provide real value.

Beyond just coding, improving AI models requires subtle skills like designing effective reinforcement learning environments or managing human expert feedback. Newman questions how close we are to recursive self-improvement by asking if AIs can automate these tasks, which rely on nuanced "taste and judgment" rather than just raw computational ability.

A major frontier for AI in science is developing 'taste'—the human ability to discern not just if a research question is solvable, but if it is genuinely interesting and impactful. Models currently struggle to differentiate an exciting result from a boring one.

The best AI models are trained on data that reflects deep, subjective qualities—not just simple criteria. This "taste" is a key differentiator, influencing everything from code generation to creative writing, and is shaped by the values of the frontier lab.

Current benchmarks focus on whether code passes tests. The future of AI evaluation must assess qualitative, human-centric aspects like 'design taste,' code maintainability, and alignment with a team's specific coding style. These are hard to measure automatically and signal a shift toward more complex, human-in-the-loop or LLM-judged evaluation frameworks.

Moving beyond solving existing problems like the Millennium Prize problems, the true test of advanced AI in mathematics will be its ability to generate novel, interesting conjectures and create new, unifying definitions. This represents a higher tier of mathematical creativity, akin to the work of the greatest mathematicians who frame the questions for others to solve.

AI models, trained on data divorced from our lived, biological experience, lack the innate aesthetic sense that almost all humans possess. This makes taste and aesthetic judgment a uniquely human and valuable contribution as AI handles more logical and computational tasks.

We perceive complex math as a pinnacle of intelligence, but for AI, it may be an easier problem than tasks we find trivial. Like chess, which computers mastered decades ago, solving major math problems might not signify human-level reasoning but rather that the domain is surprisingly susceptible to computational approaches.