UK AI Safety Institute · open-source
Inspect AI
model-evalautomated-redteam
What it does
Framework for large-scale model evaluation with first-class support for agent and tool-use scenarios. Built by the UK AI Safety Institute; designed around the kinds of evaluations institutional safety teams actually need.
What it doesn't do
Aimed at sophisticated evaluators; the on-ramp is steeper than a YAML-config tool. Limited turn-key reporting; you assemble what you need.
Best for
- Teams running agentic evaluations with tool-use, multi-step reasoning
- Safety-focused organizations matching public benchmark methodology
- Researchers building reusable eval suites
Not for
- Quick proof-of-concept for non-specialists
Pricing
Free (Apache 2.0).
Integration effort
high
Team skill required
Strong Python; evaluation-method literacy; LLM-internals familiarity.
Last reviewed: 2026-05-01