Theoretical alignment research (heuristic arguments, low-probability estimation, eliciting latent knowledge) led by Paul Christiano until 2024.
ARC works on the hard mathematical problem of proving that AI systems will behave as intended. Its testing arm became the independent evaluator METR.
Focus
- Alignment research
Funding, as far as we know
Coefficient Giving (formerly Open Philanthropy) - https://en.wikipedia.org/wiki/Coefficient_Giving
What they've actually done
Spun out its evaluations team as METR (Model Evaluation and Threat Research), which now does pre-deployment testing for labs
2023-12
Published a 'bird's eye view' of its research agenda on formal heuristic arguments and low-probability estimation
2024-2025
Founder Paul Christiano left to become head of AI safety at the US AI Safety Institute (April 2024), later leaving government
2024-04