Research on 'AI control': keeping AI safe even if models are misaligned, plus empirical alignment studies.
Redwood studies how to keep AI systems in check even if they secretly have the wrong goals. Its 2024 work with Anthropic showed a model pretending to go along with training.
Focus
- Alignment research
- Safety evaluations
- Human oversight
Funding, as far as we know
Open Philanthropy grants (multi-million, 2021-2023) - https://www.openphilanthropy.org/grants/?q=redwood
What they've actually done
Co-authored with Anthropic the 'Alignment Faking in Large Language Models' paper, first empirical evidence of a model strategically faking compliance
18 December 2024
Developed the 'AI control' research agenda (trusted monitoring, untrusted-model protocols) adopted by lab safety teams
2023-2025