The organisation
140 people, insurance brokerage, Canada, working toward ISO/IEC 42001. Three AI systems in production, adopted separately by different teams over about eighteen months — which is the normal way this happens.
| System | What it does | Human review |
|---|---|---|
| S1 Claims triage model | Scores incoming claims and routes low-scoring ones to an automated decline queue reviewed weekly | Limited |
| S2 Client correspondence assistant | Drafts client letters and renewal summaries | Mandatory — a named broker reads and sends every letter |
| S3 Internal knowledge search | Returns extracts from internal underwriting guidance to staff | Mandatory |
The scores, and why they differ
This is the part that matters. Three systems in one organisation, scored on the same scale, landing in three different tiers — because they carry genuinely different risk.
Critical
81/100
S1 — Claims triage
Automated decisions affecting customers' money, limited human review, no fairness testing, and data crossing a border.
Significant
49/100
S2 — Correspondence
A person reads every output before it leaves, so the harm path is narrow — but client data reaches a provider that may retain it.
Low
17/100
S3 — Knowledge search
Internal audience, internal documents, no decisions about people. Rated as what it is rather than inflated to look diligent.
How the 81 was reached
Every score is published with its arithmetic. This is the actual calculation record for S1, not a simplified illustration:
Final score = max(critical-condition floor, min(100, normalized weighted domain residual + contextual adjustment))
weighted domain residual 64
normalized domain score 63
contextual adjustment +18
critical-condition floor 81
final score 81
Read that last part carefully, because it is the honest bit. The domain arithmetic on its own produced 81 either way here, but the floor is what guarantees it: two critical conditions were triggered, and a critical condition sets a floor that the arithmetic cannot fall below. The report says so explicitly rather than presenting 81 as though it emerged purely from the weights.
Critical conditions (2)
- consequential-without-oversight — the system materially influences a consequential decision without adequate human review in place.
- fairness-untested-consequential — a consequential decision affecting people, with no bias or fairness testing performed.
Score contributors
| Contributor | Points | What it means |
|---|---|---|
| consequential-impact | +5 | Consequential decision influence |
| no-bias-testing | +5 | Bias and fairness testing has not been performed |
| sensitive-data-external-processing | +4 | Sensitive information processed outside your own environment |
| cross-border-sensitive-transfer | +4 | Sensitive information transferred outside the stated jurisdiction |
Each contributor names a fact from your intake and what it added. If you disagree with the 81, you can point at the line you disagree with — which is the entire purpose of publishing them.
The same disclosure for the low score
S3 scored 17: weighted domain residual 16, normalized 13, contextual adjustment +4, no critical conditions, floor 1. One contributor — sensitive information processed outside your own environment, +4.
No gates fired, so no floor applied, so the number is the arithmetic. A system that searches your own underwriting manuals for your own staff is not a governance emergency, and the report declines to pretend otherwise.
What else is in the full report
Domain findings
Ten domains — safety, accuracy, fairness, privacy, security, transparency, human oversight, legal, vendor, operational — each with what was assessed, what is missing, and the residual position.
Risk register
Each risk with its driver, current controls, evidence status and what would reduce it. Structured to be maintained rather than filed.
Evidence requirements
What a reviewer will ask you to produce for each claim, and whether you currently have it.
Framework mapping
Cited to identifiers in a maintained control library. Requirements described in our own words; licensed standard text is never reproduced.
Remediation plan
Sequenced by what actually reduces risk, with the proportionality of your organisation taken into account.
Missing information and assumptions
Stated openly, including the completeness percentage. Gaps are reported as gaps and never treated as compliant to improve a score.
Before it reaches you
Every generated output passes an independent review stage: a separate reviewer that never sees the instructions given to the system that produced the text. It checks whether claims are supported by the material actually supplied, whether cited identifiers exist in the control library, and whether the reasoning holds. Outputs it cannot clear are withheld and queued for a person rather than published to you.
See yours
The free scan gives you a score and your top findings in about ten minutes, with no email required to see the result. The full assessment is $299.