Evaluate frontier models for deceptive and 'scheming' behaviour and build a science of scheming.
Apollo is a London lab that tests whether AI models will lie, sandbag, or try to evade shutdown. Its findings appear in the labs' own model reports.
Focus
- Safety evaluations
- Alignment research
Funding, as far as we know
Philanthropic grants (incl. Open Philanthropy / Coefficient Giving); specifics beyond that unknown
What they've actually done
Published 'Frontier Models are Capable of In-context Scheming', showing o1, Claude and others could sabotage oversight in tests
5 December 2024
Collaborated with OpenAI on 'deliberative alignment' anti-scheming evaluations and training
17 September 2025
Conducted pre-deployment scheming evaluations cited in OpenAI, Anthropic and Google DeepMind system cards
2024-2026