BlogReference

ISO 42001 certification — what the external audit actually checks

Internal audit readiness is not the same exercise as the certification audit itself. Here is what an ISO/IEC 42001 external auditor actually samples, the nonconformity categories that decide pass or fail, and the evidence gaps that resurface most often at Stage 2.

Drel11 min read

Preparing for an ISO/IEC 42001 internal audit and preparing for the external certification audit feel like the same exercise from the inside — the same evidence checklist, the same control walkthrough, the same people in the same room. They are not the same exercise, and the difference matters more than most organisations expect the first time an accredited certification body shows up.

The internal audit is run by your own organisation (or a consultant acting on its behalf), against your own interpretation of the standard, with the freedom to fix what it finds before anyone external sees it. The certification audit is run by an accredited third party whose job is specifically to decide whether your AI management system (AIMS) meets the standard well enough to issue a certificate — and whose finding, unlike an internal audit finding, is not something you get to quietly correct before it counts.

This piece is about what that external audit actually checks: the two-stage structure, how the auditor samples evidence, what counts as a nonconformity versus an observation, and the specific gaps that surface at Stage 2 even in organisations whose internal audit found nothing.

Why this is not the internal audit again

The internal audit (covered in iso-42001-internal-audit) exists to find and fix problems before an external party sees them. It is formative — its value is in the fixing, not the finding. The certification audit is summative: its job is to produce a judgement about whether your AIMS conforms, and that judgement becomes the basis for a certificate other parties — customers, regulators, procurement reviewers — will rely on.

This distinction changes what “ready” means. An organisation can pass its own internal audit cleanly and still be unprepared for the certification audit, because the internal audit was run by people who already knew where the AIMS was weak and unconsciously steered around it, or because the internal audit sampled the same three well-documented controls every time rather than sampling broadly the way an external auditor will.

Stage 1 — documentation review

Stage 1 is a readiness check, not a full audit. The certification body reviews your documented AIMS — scope statement, policy, risk assessment methodology, Statement of Applicability, the Annex A controls you have selected and how you have justified excluding any you have not — against the structure the standard requires. The auditor is checking whether the documentation is complete and internally consistent enough to make a Stage 2 audit meaningful, not whether the AIMS is actually working yet.

Stage 1 typically includes a site or remote visit and interviews with key roles to confirm the organisation understands its own documented system — not just that the documents exist, but that the people responsible for them can describe how the system actually operates. A common Stage 1 finding is a scope statement that does not match reality: it names AI systems the audit later discovers are out of date, or omits a system that should clearly be in scope. Fixing scope-statement gaps at Stage 1 is cheap. Discovering them at Stage 2 is not.

Stage 1 vs Stage 2 — what changes between them

Stage 1Stage 2
What it checksDoes the AIMS documentation exist and cover the required scope?Is the AIMS actually operating as documented, with evidence?
Primary methodDocument review, readiness interviewSampling: interviews, records, system walkthroughs
Typical timingWeeks before Stage 2After Stage 1 findings are closed
Consequence of failureStage 2 is postponed until documentation gaps closeCertification is withheld until major nonconformities are closed

Stage 2 — implementation audit

Stage 2 is where the certification decision actually gets made. The auditor is now checking whether the AIMS documented at Stage 1 is operating in practice: are risk assessments actually being performed on the cadence the policy describes, are the Annex A controls you claimed as applicable actually implemented and evidenced, does management review actually happen with the inputs the standard requires, is the internal audit programme itself functioning.

This stage is evidence-heavy in a specific way that catches organisations off guard: the auditor is not asking “do you have a policy that says you do X” — Stage 1 already confirmed the policy exists. Stage 2 asks “show me the record that X actually happened, for a system or time period I choose, not one you choose.” An organisation whose evidence exists only for the systems it expected to be asked about, and not for others in scope, fails this stage even though its documentation was fine.

Stage 1 asks whether the map exists. Stage 2 asks whether the territory matches the map — for whichever part of the territory the auditor decides to walk.

How the auditor samples

Certification audits are sampling exercises, not exhaustive reviews — no auditor reviews every AI system, every control, every record in the time allotted. What the auditor samples, and how, is the part organisations prepare for least, because internal audits are often designed around demonstrating specific, chosen examples rather than surviving an externally chosen sample.

A competent auditor samples across at least three dimensions: across systems (not just the flagship AI system the organisation is proudest of), across time (records from several months, not just the most recent), and across roles (talking to the people who actually do the work, not only the AIMS owner who wrote the documentation). A gap that only shows up for a system the organisation did not expect to be sampled, or a time period before the AIMS owner tightened a process, is exactly the kind of gap Stage 2 sampling is designed to surface.

Nonconformity categories

Findings at Stage 2 are not binary pass/fail per control — they are classified, and the classification determines what happens next.

  • Major nonconformity. A systemic failure, or an absence of a required element of the AIMS entirely — for example, no evidence risk assessments have been performed for a system clearly in scope. Certification is withheld until the major nonconformity is corrected and the correction is verified, typically requiring a follow-up visit.
  • Minor nonconformity. An isolated lapse in an otherwise functioning process — a single missed management review meeting, one system's risk register slightly out of date. Certification can typically proceed with a corrective action plan and a deadline, verified at the next surveillance audit rather than immediately.
  • Observation / opportunity for improvement. Not a nonconformity — the process meets the standard, but the auditor notes a weakness likely to become a problem if left unaddressed. No corrective action is required, though ignoring repeated observations across audit cycles tends to produce an actual nonconformity later.

The distinction between major and minor is frequently about breadth and system-ness rather than severity in the security sense. A single missed record is minor. The same gap repeated across every system in scope, or absent for the process entirely, is major — because it indicates the control is not actually operating, not merely that one instance of it slipped.

Findings that recur at Stage 2

Across certification cycles we have seen or reviewed evidence for, the same handful of gaps show up disproportionately at Stage 2 — not because they are exotic, but because they are the parts of an AIMS that are easy to document and easy to neglect operationally.

  • Risk assessments that exist but are stale. The methodology document is solid; the actual risk assessments for specific systems were performed once, at AIMS launch, and never refreshed against the standard's expected cadence or the organisation's own re-assessment triggers.
  • Management review that is a meeting, not an input-driven process. Minutes exist, but they do not show the specific inputs the standard requires — audit results, risk assessment outcomes, nonconformity status, changes in context — being actually discussed and acted on.
  • Scope drift between the Statement of Applicability and reality. New AI systems entered production after the AIMS was documented and were never formally brought into scope, so their controls and evidence do not exist even though the systems clearly meet the scope criteria.
  • Internal audit programme that only ever finds minor issues. An internal audit programme that has never once raised a significant finding reads, to an external auditor, as either an unusually mature AIMS or an internal audit that is not actually independent and rigorous. The latter is the more common explanation, and it invites closer external scrutiny rather than less.
  • Annex A control justifications that do not hold up under questioning. A control marked “not applicable” in the Statement of Applicability, with a justification that sounded reasonable on paper but does not survive the auditor asking a follow-up question about the specific system it was excluded for.

After certification: surveillance audits

Certification is not a one-time event. Accredited certificates typically run on a three-year cycle with annual surveillance audits in between — narrower than the full Stage 2 audit, but real, and specifically designed to check that corrective actions from the previous audit were actually closed and that the AIMS has kept pace with organisational change: new systems, new regulatory obligations, structural changes to the governance committee.

Organisations that treat certification as a project with an end date, rather than an operating discipline with a recurring external check, tend to arrive at the first surveillance audit with the same gaps Stage 2 would have found a year earlier if the AIMS had not been left on autopilot after the certificate was issued.

What to prepare that internal audit doesn't cover

Beyond the standard evidence checklist, three things specifically improve Stage 2 readiness because they address how an external, sampling-based audit works rather than how an internal, self-directed one does:

  • Confirm evidence exists for every system currently in scope, not just the systems the internal audit historically sampled.
  • Pull the last 6–12 months of records for each recurring process (risk assessment cycle, management review, internal audit) and confirm they read as a continuous operating history, not a set of documents created just before the audit.
  • Have the people who actually do the work — not only the AIMS owner — ready to be interviewed and able to describe the process in their own words, consistent with the documentation.
  • Re-check every Annex A exclusion in the Statement of Applicability against current systems, and be ready to justify each one against a specific, named system rather than a general statement.

None of this replaces the internal audit — it extends it, by preparing specifically for the parts of Stage 2 that a self-directed internal review tends not to probe as hard as an external, accredited one will.

Keep certification evidence current between audits.

Drel produces the control plan, risk register, and evidence pack artefacts an ISO/IEC 42001 AIMS needs — kept current per system, so Stage 2 sampling finds a continuous operating record instead of documents assembled just before the auditor arrives.

A note on scope: Drel reviews assessed systems against documented architecture, configuration and intent. It does not ingest live telemetry from production environments. Dispositions reflect the assessed system at the time of review and the re-assessment triggers that govern when the disposition must be revisited.