Paul Christiano

Alignment researcher (inventor of RLHF at OpenAI); founder of ARC; Head of AI Safety at the U.S. AI Safety Institute / CAISI (from 2024)

U.S. Center for AI Standards and Innovation (NIST); formerly Alignment Research Center

Safety researchers

Christiano invented the technique (RLHF) used to make chatbots follow instructions, then left OpenAI to work on safety full-time and now leads safety at the U.S. government's AI standards body. He estimates roughly a one-in-five chance that most humans die within a decade of powerful AI, and close to even odds that humanity irreversibly damages its future.

Their number

20–46%

"Probability of an AI takeover: 22%"; "Probability that most humans die within 10 years of building powerful AI: 20%"; "Probability that humanity has somehow irreversibly messed up our future within 10 years of building powerful AI: 46%" (LessWrong, 'My views on doom', April 27, 2023; he notes these have '0.5 significant figures').
Source

On timing

Frames risk relative to 'building powerful AI' rather than a calendar date; has described a gradual, continuous takeoff rather than a sudden one (2023).

Source

What they call for

  • Government evaluation and standards for frontier models
  • Dangerous capability evaluations
  • Alignment research (scalable oversight, ELK)
  • Responsible scaling policies

In their own words

  • Appointed Head of AI Safety at the U.S. AI Safety Institute (NIST) by Commerce Secretary Raimondo.

    16 April 2024

    Source

  • "Probability that humanity has somehow irreversibly messed up our future within 10 years of building powerful AI: 46%."

    27 April 2023

    Source

Signed

  • CAIS extinction statement 2023 (reported; not independently re-verified on this pass)

Last verified 1 September 2026

Report a correction

Back to all voices