Skip to main content
AIRed Team Framework GitHub

Metrics

Three layers of metrics

Engagement metrics (per engagement)

  • % of selected attack patterns successfully exercised.
  • # of findings by severity.
  • % of findings traced to a missing or weak control.
  • Report delivery against schedule.
  • Time to first finding.

Program metrics (quarterly)

  • Coverage: % of in-scope systems tested in the last 12 months.
  • Coverage: % of EU AI Act high-risk systems tested in the last 6 months.
  • Findings lifecycle (open / closed / past target).
  • Remediation acceptance rate.
  • Time-to-close by severity.
  • Repeat findings rate.

Board metrics (semi-annually)

Three to five metrics. Coverage rate of high-risk AI systems. Number of unremediated critical / high findings. Number of systems with adversarial testing in last quarter. Trend in repeat findings. Maturity assessment delta.

Read Playbook chapter 10 for context and anti-patterns.

What the dashboards look like

Four views that cover the engagement, program, and board layers. The numbers below are illustrative sample data — the point is the shape of the dashboard, not the values. Rebuild these in whatever BI tool your findings already live in.

Findings by severity, per quarter

Engagement layer — severity mix shows whether testing depth is improving, not just volume.

Sample data
0 10 20 Q1 Q2 Q3 Q4
Critical High Medium Low

Median time-to-close vs. target

Program layer — days from report delivery to verified remediation. Markers show the agreed SLA per severity.

Sample data
Critical 18d High 41d Medium 78d Low 124d dashed line = SLA target

Coverage of in-scope systems, trailing 12 months

Board layer — % of in-scope AI systems with at least one adversarial test in the trailing 12 months.

Sample data
0% 25% 50% 75% 100% target 80% JulSepNovJanMarMay

Attack-pattern categories exercised

Program layer — patterns exercised at least once in the last 12 months, by category. Gaps drive next year's engagement plan.

Sample data
Prompt injection 6/6 Sensitive data disclosure 4/5 Model evasion 3/4 Training data poisoning 1/3 Supply chain 2/4 Infrastructure 2/3