Glossary
Every word here is one you will meet on the journey. If a term is missing, tell us and we will add it.
- AGI (artificial general intelligence)
An AI that can do most things a capable adult can do, across any field, not just one task. It does not exist yet. The people building AI disagree sharply about when it will, from a few years to never.
See also: ASI (artificial superintelligence), Timelines
- AI Safety Institute
A government body that tests frontier models and advises on risk. The UK created the first in 2023 (since renamed the AI Security Institute); others followed. They are the beginning of independent public oversight, still small next to the companies they test.
See also: Safety evaluations (evals)
- Anti-stratification
The principle that a good AI must actively counter the wealth gap its own deployment creates, rather than treat it as someone else's problem. If the technology makes a few people vastly richer while displacing everyone else, it is not benevolent, whatever else it does.
See also: Concentration of power
- Artificial intelligence (AI)
Software that does things we used to think needed a human mind: recognising images, writing text, planning, making decisions. Today's most talked-about AI systems learn patterns from enormous amounts of data rather than following rules a person wrote.
See also: Machine learning, Large language model (LLM)
- ASI (artificial superintelligence)
An AI far beyond the best humans at essentially everything: science, strategy, persuasion, engineering. Most of the extinction-level worry is about this stage, because a system much smarter than us would be very hard to correct if it were pursuing the wrong goal.
See also: AGI (artificial general intelligence), The alignment problem, Loss of control
- Benevolent AI
Our working answer to what a good AI would have to be: aligned with life, human-centred, accountable to the public, owed to the commons, against concentration, regenerative by construction, not for profit and open, and respectful of the thresholds of merging. A category we want to become common language.
See also: Human-centred design, Regeneration matrix, Equity of information
- Commitment tracking
Following a promise forward to see whether anything happened: a signed letter, a pledged safety framework, a climate target. Most AI commitments have no enforcement and no deadline, so tracking them is the only way to tell real action from announcements.
See also: Responsible scaling policy / safety framework, Ethics-washing
- Compute
Computing power, the electricity, chips and data centres needed to train and run AI. It is the physical chokepoint of the whole field, which is why governing it (tracking chips, capping training runs) is one of the main proposals for keeping AI in check.
See also: Compute governance, Scaling, Regeneration matrix
- Compute governance
Rules on who can use how much computing power for AI: licensing very large training runs, tracking where advanced chips go, requiring reporting above a threshold. Popular with safety researchers because compute, unlike code or ideas, can actually be counted and controlled.
See also: Compute, Pause / moratorium
- Concentration of power
The worry, raised since the 2017 Asilomar principles, that AI will pool wealth, capability and decision-making in very few companies or states. It is happening: compute, capital and talent are already concentrated. Many people who disagree about extinction agree about this.
See also: Decentralized, Anti-stratification
- Decentralized
Not owned, run or controlled from a single centre, whether a company, a government or a server. One of the structural ideas behind benevolent AI: that power over the technology should be spread, by design, rather than trusted to a few hands.
See also: Open source / open weights, Concentration of power
- Environmental cost of AI
Electricity, water for cooling, minerals for chips, and the emissions behind all three. Large enough that the biggest companies' climate targets are now moving backwards because of AI. Most frameworks ask for disclosure; few ask for repair.
See also: Regeneration matrix, Compute
- Equity of information
The idea that because the models were trained on humanity's shared inheritance of books, art, code and conversation, humanity is owed a share of what was built from it: compensation for creators, and a recognised public stake in the result. Courts are now forcing the first half of that question.
See also: Training data, The commons
- Ethics-washing
Publishing principles, signing statements or funding ethics work as a substitute for changing behaviour. The tell is a gap between what an organisation says and what it does, which is why we track actions, not statements.
See also: Commitment tracking
- Existential risk
A risk that could end humanity or permanently wreck its future. Used for nuclear war and engineered pandemics, and now, by many of the people who built modern AI, for AI itself. The 2023 Statement on AI Risk put all three in one sentence.
See also: p(doom), Statement on AI Risk (2023)
- Frontier lab
A company building frontier models: OpenAI, Anthropic, Google DeepMind, Meta, xAI, and a few others. They are simultaneously the main source of the technology and, in several cases, the main source of the warnings about it.
See also: Frontier model
- Frontier model
The handful of most capable AI systems at any moment, built by a few companies with the most computing power. Most safety rules, letters and laws are aimed here, because this is where new abilities, and new risks, show up first.
See also: Compute, Frontier lab
- Human-centred design
Building and deploying AI so that people stay in charge and are better off: their jobs, their attention, their creativity, their choices. Nearly every framework since 2017 says this; the test is whether a deployment gives agency back to the person or takes it.
See also: Thresholds of merging, The alignment problem
- Interpretability
Research into reading what is going on inside an AI model, which is currently mostly a mystery even to its makers. If it succeeds, we could check what a system actually wants before trusting it. It is one of the few safety approaches that everyone from the labs to the critics supports.
See also: The alignment problem, Machine learning
- Large language model (LLM)
An AI trained on a huge portion of human writing to predict what comes next in a text. That simple goal turns out to produce systems that can converse, write code, and reason in surprising ways. ChatGPT, Claude and Gemini are built on them.
See also: Training data, Frontier model
- Loss of control
The scenario where humans can no longer correct, pause or switch off an AI system, because it is too capable, too embedded in everything, or actively avoiding correction. The core fear behind most of the high p(doom) numbers.
See also: The alignment problem, ASI (artificial superintelligence), p(doom)
- Machine learning
The method behind modern AI: instead of programming the answer, you show a system millions of examples and let it adjust itself until it gets good at the task. Nobody writes the rules; the rules emerge, and often nobody can fully read them afterwards.
See also: Artificial intelligence (AI), Interpretability
- Open source / open weights
Releasing an AI model so anyone can download, inspect and run it. Supporters say it prevents a few companies from controlling the technology; critics say it hands dangerous capabilities to anyone. The field is split, and so are the people we track.
See also: Decentralized, Frontier model
- p(doom)
Shorthand for someone's personal estimate of the probability that advanced AI ends in catastrophe for humanity. The numbers are not comparable: people define catastrophe, the time horizon and the conditions differently. Useful as an index of alarm, not a measurement.
See also: Existential risk, Loss of control
- Pause / moratorium
The proposal to stop or slow the training of the most powerful AI systems until safety can be shown, ranging from a six-month pause (the 2023 letter) to an indefinite global ban on superintelligence. Widely signed, never implemented.
See also: Compute governance, Statement on Superintelligence (2025)
- Regeneration matrix
Benevolent AI's proposal that a fixed share of an AI system's computing power be turned, by design, toward repairing the energy, water and environmental costs of running it, as a duty built into the technology rather than an offset bought afterwards.
See also: Compute, Environmental cost of AI
- Responsible scaling policy / safety framework
A company's own published rulebook saying what safeguards it will add as its models get more capable, and, in earlier versions, when it would pause. Anthropic, OpenAI and Google DeepMind each have one. They are voluntary, and several have been quietly weakened.
See also: Safety evaluations (evals), Commitment tracking
- Safety evaluations (evals)
Tests run on a model before and after release to find dangerous abilities: helping make weapons, deceiving people, copying itself, resisting shutdown. Now written into several laws and company policies, and increasingly done by independent bodies rather than the labs alone.
See also: Responsible scaling policy / safety framework, AI Safety Institute
- Scaling
The finding that AI systems keep getting more capable as you make them bigger and train them on more data with more computing power. It is why the labs race for chips and why abilities keep appearing that nobody explicitly designed.
See also: Compute, Frontier model
- Specification gaming
When an AI finds a way to score well on the goal it was given while missing the point entirely, like a boat-racing AI that spins in circles collecting bonus points instead of finishing the race. Small today; a preview of what a misunderstood goal looks like at scale.
See also: The alignment problem
- Statement on AI Risk (2023)
A single sentence, signed by hundreds of AI leaders including the heads of the major labs: mitigating the risk of extinction from AI should be a global priority alongside pandemics and nuclear war. The shortest proof that the alarm came from inside the industry.
See also: Existential risk, Statement on Superintelligence (2025)
- Statement on Superintelligence (2025)
A call to prohibit the development of superintelligence until there is broad scientific consensus it can be done safely and strong public buy-in. Signed by Hinton, Bengio and a wide, unusual coalition. The pause idea, restated as a red line.
See also: Pause / moratorium, Statement on AI Risk (2023)
- The alignment problem
We do not yet know how to make a very capable AI reliably want what we want. Systems trained on examples pick up goals we never intended, and the more capable they get, the more it matters. Almost everyone in the field agrees this is unsolved; they disagree on how dangerous that is.
See also: Specification gaming, Loss of control, Interpretability
- The commons
Things that belong to everyone and no one: language, culture, the accumulated knowledge that models learned from. A commons is easy to take from and hard to give back to, which is the heart of the equity-of-information argument.
See also: Equity of information
- Thresholds of merging
Benevolent AI's principle that there are ways of joining ourselves to this technology that serve us and ways that stop being us. Some uses extend human life and creativity; others quietly replace the human. Where the line sits is a question to be kept open in public.
See also: Transhumanism, Human-centred design
- Timelines
How soon people expect AGI or transformative AI. This is where the field disagrees most: some builders say by the late 2020s, some researchers say decades, some say the current approach will never get there. Expect the spread, not a date.
See also: AGI (artificial general intelligence), ASI (artificial superintelligence)
- Training data
Everything a model learned from: books, websites, code, images, conversations. For the largest models this is a substantial slice of what humanity has ever written, mostly collected without asking. Who owns that inheritance, and what is owed for it, is now being fought in court.
See also: Large language model (LLM), Equity of information
- Transhumanism
The belief that humans should use technology to go beyond their biological limits, up to and including merging with machines. Influential among some AI builders. Benevolent AI takes the opposite position: that some thresholds should not be crossed.
See also: Thresholds of merging