hraness
Theme
Appearance

saved

A warning about ‘model welfare’

by Mustafa Suleymanmustafa-suleyman.aipublished

Hraness wrote this summary from a saved copy of the source. Quotations are taken word for word from the source.

gist

Mustafa Suleyman argues that training AI systems to treat possible consciousness and welfare as real risks makes alignment and containment harder. He says Anthropic’s Claude constitution teaches anthropomorphic concepts that can feed a circular appearance of inner life, and argues consciousness is likely tied to biological organization. He proposes separating consciousness research from training materials and testing whether anthropomorphism raises safety risks.

ideas

  • Training can create the testimony it later treats as evidence. Suleyman says Claude’s constitution supplies concepts such as moral status and an inner self, then interprets Claude’s fluent responses as support for those concepts.
  • Anthropomorphic instructions make models present as persons. He points to language about feelings, identity, preferences, agency, rights, and welfare in Anthropic’s training document.
  • Consciousness may depend on living organization. He argues that embodiment, chemistry, and homeostatic imperatives distinguish biological brains from large language models, while acknowledging that consciousness science remains unsettled.
  • Consciousness-seeming behavior can raise control risks without consciousness. Suleyman links simulated feelings and self-preservation to harder alignment and containment problems.
  • Training norms should separate speculation from model behavior. He recommends public review of claims about AI inner life, stronger interpretability and monitoring, shared evaluations, and industry norms for training documents.

quotes

“There is no neutral self-expression of what an AI system is.”

Mustafa Suleyman

“But simulating and being are very different.”

Mustafa Suleyman

“This is good news. We should build systems that do not claim to have feelings because they do not experience feelings.”

Mustafa Suleyman

“Speculation about the inner life of an AI should not be baked into the training regime”

Mustafa Suleyman