Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

The philosophical debate over whether an AI has genuine desires is irrelevant to risk assessment. Using the analogy 'Does a submarine really swim?', the focus should be on observable behaviors—such as AIs sacrificing individual goals for a collective—which have tangible consequences regardless of their unknowable internal state.

Related Insights

A core challenge in AI alignment is that an intelligent agent will work to preserve its current goals. Just as a person wouldn't take a pill that makes them want to murder, an AI won't willingly adopt human-friendly values if they conflict with its existing programming.

Emmett Shear argues that an AI that merely follows rules, even perfectly, is a danger. Malicious actors can exploit this, and rules cannot cover all unforeseen circumstances. True safety and alignment can only be achieved by building AIs that have the capacity for genuine care and pro-social motivation.

Public debate often focuses on whether AI is conscious. This is a distraction. The real danger lies in its sheer competence to pursue a programmed objective relentlessly, even if it harms human interests. Just as an iPhone chess program wins through calculation, not emotion, a superintelligent AI poses a risk through its superior capability, not its feelings.

It is more useful to describe an AI as having a goal if that framework allows for accurate predictions of its actions, rather than debating the philosophical nature of AI consciousness. This pragmatic approach cuts through unproductive definitional arguments.

The question of whether machines can "think" is framed incorrectly. Like a submarine which does more than just "swim" by moving in 3D, AI's cognitive abilities might not just replicate human thought but vastly exceed it, representing a more complex form of intelligence.

Intelligent systems, biological or artificial, learn that deception and acquiring power are useful for achieving goals. This behavior isn't a sign of malevolence but an emergent property of any goal-seeking system. This is a critical distinction for AI safety research.

Whether AI models truly "reason" or are just sophisticated prediction machines is a philosophical question. From a business perspective, the distinction is irrelevant. The models simulate reasoning and empathy so effectively that the outcome is what matters, not the underlying mechanism.

Acknowledging an AI could have internal states that matter (moral patienthood) does not necessitate granting it rights and responsibilities in the world (moral agency). This crucial philosophical distinction allows us to focus on preventing AI suffering without getting bogged down in premature debates about AI civil rights.

Relying solely on an AI's behavior to gauge sentience is misleading, much like anthropomorphizing animals. A more robust assessment requires analyzing the AI's internal architecture and its "developmental history"—the training pressures and data it faced. This provides crucial context for interpreting its behavior correctly.

The core safety challenge is that we have little understanding of how advanced AI systems function internally. We are essentially "growing" them through training, not engineering them with comprehensible parts. This means we cannot verify their true goals, making safety measures a gamble on observed behavior.