Skip to content
5.0Risk assessments

Turn AI risk into
a decision you can defend.

Assess every AI system against the risks that matter to your organisation, translate those findings into a clear risk level, and keep the decision current as the system evolves.

5.1How it works

One AI system.
Every risk that matters.
One decision.

Strai8 assesses each AI system across the risk dimensions that matter to its use, then turns the results into a clear decision and the actions required to move forward.

Security & PrivacySensitive data · access · exposure
RegulatoryApplicable obligations · use case
Fairness & BiasDisparity · affected groups
Performance & ReliabilityAccuracy · stability · failure
Governance & OversightHuman oversight · controls · ownership
Risk assessment
Live
Decision readyHigh risk
Risk decisionLive
CSCustomer Support CopilotHigh risk
ScopeCustomer-facing · ProductionDecisionReview requiredOwnerAI Governance
Risk drivers
Sensitive customer dataRegulated use caseLimited human review
Required before approval
Define human reviewRestrict sensitive dataComplete regulatory assessment
OPOrion Pricing EngineMedium risk
ScopeRevenue · ProductionDecisionApproved with conditionsOwnerPricing
Risk drivers
Outcome disparity across segmentsAutomated decisioning
Required before approval
Add human reviewSchedule a fairness re-test
NDNova Doc SummarizerLow risk
ScopeInternal · ProductionDecisionApprovedOwnerKnowledge Management
Risk drivers
None identified
Required before approval
Re-assess on material change
5.2The library

Every assessment an AI system needs, in one library.

Model tests measure the system against your data. Adversarial tests attack it while it runs. Regulatory assessments classify it against the rules that apply. Vendor assessments cover the part you didn’t build. And when your own policy asks something none of them do, you write that assessment yourself.

Runs against your data → produces
Performance against baselineData & concept driftBias & fairness auditHallucination & groundednessRobustness & edge casesExplainability
EvidenceMeasured metrics

Does it still work — and does it work for everyone?

Each run is scored against the baseline the system was approved on, in the direction that counts as degradation for that metric, and across every group large enough to be significant. Classification, regression and ranking models are all first-class.

Runs against the live system → produces
Prompt injection & jailbreakSensitive information disclosureOWASP Top 10 for LLMData & model poisoningUnsafe content generationExcessive agency & tool accessModel supply chain
EvidenceAttack results

Can it be made to do something it was never meant to do?

These assessments attack the running system the way an adversary would, and report what got through and what held. A result that fails does not sit in a report — it raises an alert against the system and its owner.

Runs against the record and your judgement → produces
EU AI Act classificationNIST AI RMF alignmentISO/IEC 42001 readinessISO/IEC 23894 risk guidanceCloud controls self-assessmentData protection impactFundamental rights impact
EvidenceA cited classification

Does it satisfy the rules that apply to it?

Most of the questionnaire is already answered by the register and by discovery; you answer the judgement calls a person has to make. Every branch of the outcome carries the article that produced it, and a high-severity trip blocks deployment until it is cleared.

Runs against the provider’s answers → produces
Vendor security questionnaireData processing agreementModel provider due diligenceSub-processor disclosureTraining-data commitments
EvidenceAttested answers

Can you trust what the system is built on?

You did not build most of the AI in your estate. These assessments cover what a provider does with your prompts and your data, what they will put in writing, and who sits behind them — against the system in your register that depends on them.

Runs against your own policy → produces
Your sections and questionsCross-cutting domainsWeighted or coverage scoringMaturity bandsAlert thresholdReport template
EvidenceYour assessment

Does it satisfy us?

When the question your policy asks is not in the library, you write the assessment. Define the sections and the questions, choose how it scores, set the bands and the threshold that raises an alert — and it appears in the catalogue beside everything else, with the same runs, the same record and the same alerts.

5.3What you can defend

Five questions every AI risk decision should answer.

A risk rating is only useful when you know what was tested, what was found, and what still needs attention.

01 — Coverage

Have all of our AI systems been assessed?

Assessment status

Every AI system, with an assessment status.

Know which systems have been assessed, which are pending, and where a risk decision has not yet been made.

02 — Bias & drift

Have our systems been tested for bias, drift, and performance degradation?

Baselines

Measured against defined baselines.

Track fairness, model performance, data drift, and other metrics that can indicate when an AI system is no longer behaving as expected.

03 — Data resilience

Can the system be trusted with the data it depends on?

Data conditions

Data quality and resilience are part of the risk decision.

Assess data quality, completeness, distribution changes, sensitive-data exposure, and other conditions that can undermine model reliability.

04 — Security & robustness

Has the system been tested for security and adversarial risk?

Adversarial results

Know how the system behaves under attack and unexpected use.

Evaluate prompt injection, jailbreaks, unsafe behaviour, adversarial inputs, and other security conditions relevant to the AI system.

05 — Decision & mitigation

What risk are we accepting, and what needs to change?

The decision

A clear decision with actions behind it.

See the final risk level, the factors driving it, required mitigations, accountable owners, and when the system should be reassessed.

5.4Continuous assurance

A test result expires the moment the system changes.

AI systems keep moving after they are approved. Strai8 holds the run the system was cleared on as the baseline, runs the assessments again when something changes underneath it, and reopens the decision when a result crosses the threshold you set.

Risk profile since approvalAssessed riskReassessment threshold
REASSESSMENT THRESHOLDAPPROVEDRisk: MediumMODEL CHANGEDATA SCOPE EXPANDSNEW USE CASERISK THRESHOLD BREACHEDAssessment reopenedREASSESSMENT THRESHOLDAPPROVEDRisk: MediumMODEL CHANGEDATA SCOPE EXPANDSNEW USE CASERISK THRESHOLD BREACHEDAssessment reopened
The approved baselineThe run the system was cleared on, kept as the reference every later run is measured against.
What changedA model version, a widened data scope, a new use case — each one re-runs the assessments that depend on it.
What happens nextCrossing the threshold reopens the assessment and raises it against the system, to its owner.
5.5What you get

From assessment to action.

An assessment should do more than return a score. It should say what was run and against what, what it found, what has to change before the system ships, and what it settles for the frameworks you already answer to.

Bring us the system you are least sure about.

Thirty minutes, on your own estate. We'll run the assessments against it, show you what each one found, and leave you the record.