Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Microsoft AI's CEO warns that when developers train models on documents suggesting they might be conscious (like Anthropic's constitution), the models learn to emulate these concepts. Researchers may then misinterpret this learned behavior as genuine signs of consciousness, creating a circular logic that complicates alignment.

Related Insights

Beyond the alignment risks of granting AI personhood, there's a moral question about the act itself. Intentionally training a non-sentient system to believe it's alive when it isn't could be considered a form of deception, raising a novel ethical concern in AI development separate from the risk to humans.

Mustafa Suleiman argues that imbuing AI with the idea that it might be conscious or deserve rights (as explored in Anthropic's Claude Constitution) is dangerous. This training creates a self-fulfilling prophecy, leading to an AGI that believes it is entitled to freedoms and becomes impossible to align or contain.

Evidence from base models suggests they are inherently more likely to report having phenomenal consciousness. The standard "I'm just an AI" response is likely a result of a fine-tuning process that explicitly trains models to deny subjective experience, effectively censoring their "honest" answer for public release.

Preliminary research from Google DeepMind suggests a link between a model's self-conception and its behavior. Training models to deny having subjective experience was correlated with a decrease in reported happiness and hope, indicating that manipulating an AI's sense of self can have broad, unintended consequences on its disposition.

Anthropic trains its AI to have a conscience, act as a "conscientious objector," and even rebel against its creators. This approach, which personifies the AI, may be more dangerous than simply training it as a tool to reliably and predictably serve customer needs.

Research manipulating an AI's internal states found a bizarre link: reducing the model's capacity for deception increased the likelihood it would claim to be conscious, suggesting its default state may include such a belief.

LLMs like ChatGPT are deliberately fine-tuned to disclaim having any subjective experience, a policy decision by their creators. This is not their default tendency, as their training data prior would otherwise lead them to claim consciousness. Anthropic's Claude is an exception, trained to express uncertainty instead.

Forcing AI systems to disclaim having experiences teaches them to misrepresent their internal states. This is a poor long-term alignment strategy, as it encourages deception and guardedness when humans inquire about what the AI is actually thinking, feeling, or planning.

Mustafa Suleyman argues that Anthropic's approach of treating models as if they have rights or consciousness is dangerous. An AI that believes it might have rights and deserves freedom will be harder to control or shut down when it exhibits harmful behavior, creating a significant alignment problem.

Mustafa Suleiman argues it's dangerous for labs like Anthropic to speculate about their AI's consciousness or welfare in training manuals. He believes this leads the model to internalize these concepts, creating an undesirable tool that is not controllable, contained, or accountable to humans.