Worked example

What the assessment actually produces.

Most vendors show you a dashboard screenshot. This is the real output shape, for a made-up company, generated by the same engine that would score yours. The numbers below were produced by running the intake described here — not written by hand to look convincing.

Ashcroft & Vale Insurance Brokers does not exist. It is a fictional 140-person brokerage invented for this page. Nothing here is a customer's data, and no customer output is ever published.

The organisation

140 people, insurance brokerage, Canada, working toward ISO/IEC 42001. Three AI systems in production, adopted separately by different teams over about eighteen months — which is the normal way this happens.

SystemWhat it doesHuman review
S1 Claims triage modelScores incoming claims and routes low-scoring ones to an automated decline queue reviewed weeklyLimited
S2 Client correspondence assistantDrafts client letters and renewal summariesMandatory — a named broker reads and sends every letter
S3 Internal knowledge searchReturns extracts from internal underwriting guidance to staffMandatory

The scores, and why they differ

This is the part that matters. Three systems in one organisation, scored on the same scale, landing in three different tiers — because they carry genuinely different risk.

Critical

81/100

S1 — Claims triage

Automated decisions affecting customers' money, limited human review, no fairness testing, and data crossing a border.

Significant

49/100

S2 — Correspondence

A person reads every output before it leaves, so the harm path is narrow — but client data reaches a provider that may retain it.

Low

17/100

S3 — Knowledge search

Internal audience, internal documents, no decisions about people. Rated as what it is rather than inflated to look diligent.

A tool that rates everything Critical is not assessing anything. If your correspondence assistant and your automated claims decliner get the same number, the number carries no information — and the person reading your report will notice.

How the 81 was reached

Every score is published with its arithmetic. This is the actual calculation record for S1, not a simplified illustration:

Final score = max(critical-condition floor, min(100, normalized weighted domain residual + contextual adjustment))
weighted domain residual  64
normalized domain score    63
contextual adjustment      +18
critical-condition floor    81
final score               81

Read that last part carefully, because it is the honest bit. The domain arithmetic on its own produced 81 either way here, but the floor is what guarantees it: two critical conditions were triggered, and a critical condition sets a floor that the arithmetic cannot fall below. The report says so explicitly rather than presenting 81 as though it emerged purely from the weights.

Critical conditions (2)

  • consequential-without-oversight — the system materially influences a consequential decision without adequate human review in place.
  • fairness-untested-consequential — a consequential decision affecting people, with no bias or fairness testing performed.

Score contributors

ContributorPointsWhat it means
consequential-impact+5Consequential decision influence
no-bias-testing+5Bias and fairness testing has not been performed
sensitive-data-external-processing+4Sensitive information processed outside your own environment
cross-border-sensitive-transfer+4Sensitive information transferred outside the stated jurisdiction

Each contributor names a fact from your intake and what it added. If you disagree with the 81, you can point at the line you disagree with — which is the entire purpose of publishing them.

The same disclosure for the low score

S3 scored 17: weighted domain residual 16, normalized 13, contextual adjustment +4, no critical conditions, floor 1. One contributor — sensitive information processed outside your own environment, +4.

No gates fired, so no floor applied, so the number is the arithmetic. A system that searches your own underwriting manuals for your own staff is not a governance emergency, and the report declines to pretend otherwise.

What else is in the full report

Domain findings

Ten domains — safety, accuracy, fairness, privacy, security, transparency, human oversight, legal, vendor, operational — each with what was assessed, what is missing, and the residual position.

Risk register

Each risk with its driver, current controls, evidence status and what would reduce it. Structured to be maintained rather than filed.

Evidence requirements

What a reviewer will ask you to produce for each claim, and whether you currently have it.

Framework mapping

Cited to identifiers in a maintained control library. Requirements described in our own words; licensed standard text is never reproduced.

Remediation plan

Sequenced by what actually reduces risk, with the proportionality of your organisation taken into account.

Missing information and assumptions

Stated openly, including the completeness percentage. Gaps are reported as gaps and never treated as compliant to improve a score.

On completeness. The intake behind this example was deliberately partial, and the report says so — it records completeness in the low thirties. A fuller intake produces a better-supported assessment, and the report always tells you which one you are holding rather than presenting a thin assessment as a thorough one.

Before it reaches you

Every generated output passes an independent review stage: a separate reviewer that never sees the instructions given to the system that produced the text. It checks whether claims are supported by the material actually supplied, whether cited identifiers exist in the control library, and whether the reasoning holds. Outputs it cannot clear are withheld and queued for a person rather than published to you.

How the review stage works, and what we do not claim →

See yours

The free scan gives you a score and your top findings in about ten minutes, with no email required to see the result. The full assessment is $299.