An ISO 42001 internal audit is the planned, independent check that your AI management system conforms to the standard and to your own stated requirements, and that it is effectively implemented and maintained. It sits in clause 9 alongside monitoring and measurement and management review, and a certification body reads its output before anything else you produced.
The mechanics transfer almost entirely from an information security audit. The difficulty does not. An ISMS audits a control environment that mostly holds still. An AI management system audits one where the model gets retrained, the vendor ships a new version, the prompt changes and the use case expands — sometimes between your fieldwork and your report.
Why an ISO 42001 Internal Audit Is Harder Than an ISMS Audit
An ISO 42001 internal audit is harder than an ISMS audit because the thing you are auditing changes while you audit it. Access control behaves the same way in March as it did in January. A model fine-tuned in January, retrained in February and swapped for a newer vendor checkpoint in March is three different systems wearing one name in the inventory.
An approach built for a stable control environment therefore produces findings that are out of date on the day they are written. You test a system, confirm the impact assessment covers it, and by the time the report is signed the use case has widened to a customer-facing surface the assessment never considered. The finding you should have raised is not "no impact assessment" — it is that nothing in the change process triggers a reassessment.
So target the processes that absorb change, not only the artefacts change leaves behind. Where the ISO 42001 requirements ask for an assessment to exist, the audit asks whether it survives the next release.
What the Internal Audit Has to Establish
The internal audit has to establish two separate things, and most internal audits answer only the first. Conformity: does the AIMS meet the organization's own requirements for it and the requirements of the standard? Effectiveness: is the system effectively implemented and maintained? Clause 9 asks for both.
Conformity is easier because it is answered with documents. The policy exists, the scope statement exists, the risk and impact assessment methods are written down. An auditor working a clause checklist can close all of it in a day and report a clean audit.
Effectiveness is the question that produces findings. It asks whether the defined process ran on the systems you actually operate, when it was meant to, and whether it changed any decision. A documented human oversight step nobody has ever used is conforming and ineffective at once. If your working papers have a column for "documented" and none for "operating on the current estate", you will produce the clean audit a certification body then contradicts.
Planning the AI Management System Audit Programme
An AI management system audit programme sets frequency, methods, responsibilities and reporting across a 12-month cycle, and it has to be based on the importance of the processes involved and the results of previous audits. Those two inputs stop it degrading into an annual sweep that treats every process as equally interesting.
Importance, for an AIMS, tracks consequence rather than complexity. A system that shapes an outcome for someone outside the organization — credit decisioning, triage, eligibility — earns more frequent coverage than an internal summarisation tool, whichever one engineering finds more interesting.
Previous results matter as much. If last year's audit found the inventory incomplete, this year's programme opens with inventory completeness rather than closing with it. Carrying weak areas forward is how a programme compounds instead of repeating.
Methods deserve an explicit decision too: oversight is answered by observation, provenance by inspection. A programme that is entirely interview-based reports what people believe their systems do.
Want an internal audit that survives the Stage 2 auditor?
Share your work email and we will review your AIMS audit programme against clause 9 and mark where your evidence would not hold on a changed estate.
Auditor Competence and Independence
Auditor competence and independence is the structural problem of an AIMS audit, because the people who understand the AI systems well enough to audit them are usually the people who built them. Independence means auditors do not audit their own work. Competence means they understand what they are looking at. In a thirty-person company those requirements point at different people.
Pretending otherwise produces one of two bad outcomes. An independent but unqualified auditor confirms documents exist and misses that the validation evidence is from a superseded model. A competent but conflicted auditor is structurally unable to find the design decision they made. Three resolutions work in practice, in rough order of cost.
- Split the estate. The ML engineer audits the systems they did not build; the platform owner audits the ones they did not deploy. Credible at small scale if the assignments are documented and genuinely disjoint.
- Pair a process auditor with a technical reader. Compliance leads and owns the conclusion; a competent person independent of that system reads the artefacts. The lead auditor's independence is what matters.
- Buy the independence. An external auditor conducting the internal audit is legitimate and common, and for a first cycle usually the cheapest route to a report the certification body will not reopen.
What does not work is the AI programme owner auditing the programme they are accountable for, then explaining that to a Stage 2 auditor. Record the competence basis and the independence basis per assignment, because ISO 42001 certification auditors ask about both. The procedural sequence is otherwise shared with an ISMS, and our ISO 27001 internal audit guide covers it in depth.
What Evidence Looks Like When You Are Auditing an AIMS
Auditing an AIMS means sampling a specific set of evidence, and the useful version of each item is narrower than the artefact.
- Impact assessments, and the decisions they changed. An assessment concluding everything is acceptable, for every system, is form-filling. Look for the one that narrowed a use case or forced an oversight step — our AI system impact assessment guide covers what a defensible one contains.
- The Statement of Applicability, against reality. Not whether it exists, but whether an excluded control is still genuinely inapplicable given what shipped. The Statement of Applicability is the document most likely to be quietly stale.
- The AI system inventory, tested for completeness. The highest-yield test in the audit, and the method is to look outside the inventory rather than inside it.
- Data provenance records. Where each dataset came from, under what terms, and whether the record was made at acquisition or reconstructed later.
- Model and version change history. What changed, when, who approved it, what was revalidated. This is what makes every other piece of evidence datable.
- Human oversight actually exercised. Designed oversight is a diagram. Exercised oversight is a log entry where a person changed an outcome.
- Third-party AI terms, and whether anyone monitors against them: what the vendor committed to on model changes, training on your data, and notice.
- Incident and concern records, including escalations that never became incidents.
The Annex A objectives tell you where to look: A.5 for impact assessment, A.6 for the life cycle, A.7 for data, A.9 for use inside the intended envelope, A.10 for third parties. Our ISO 42001 controls reference maps each objective to its evidence.
The Questions That Find Real Findings
The questions that find real findings are specific, answerable and uncomfortable. Four of them do most of the work.
Is the AI inventory complete, or does it list the systems someone remembered? Test it from outside: SaaS spend, API keys, model-provider billing, browser extensions, features marketing announced this year. Every system you find that is not listed is a finding about the discovery process.
When did the Statement of Applicability last change, and did the estate change in the meantime? Two dates settle it. If the SoA has not moved in 12 months and 3 AI features shipped, the control selection was made against a different organization than the one you have.
Can you show a case where human oversight overrode an AI output? If the answer is no for the whole period, either oversight is not being exercised or the system is never wrong. One of those is a finding; the other is not true.
What happened the last time a vendor updated a model you depend on? Ask for the notice, the assessment, the revalidation and the decision. A shrug here means the change process does not extend to changes you do not control.
Sampling Across an Estate That Changes
Sampling across a changing estate means dating every sample, because a sample from an AI estate is a sample of a point in time. Record the model version, the dataset version and the configuration in force when you tested. Without those, the finding cannot be defended later.
Audit timing relative to model changes decides what you can see. Test a system 2 weeks after a retrain and the change process is fresh while monitoring evidence for the new version barely exists. Test it after a long stable period and the monitoring looks good while you learn nothing about how change is handled. Neither window is wrong; choosing one unconsciously is.
Two habits fix most of it. Stratify so the sample includes a system that changed materially in the period and one that did not — the contrast is where change-control findings come from. And re-confirm the population at the end of fieldwork, because systems that entered scope mid-audit are the likeliest to have skipped the process.
Nonconformities, Correction and Corrective Action
A nonconformity needs both correction and corrective action, which clause 10 keeps deliberately distinct, and conflating them is how findings recur. Correction fixes the instance: write the missing impact assessment, update the inventory entry, revalidate the model. Corrective action addresses the cause so the instance does not happen again.
Root cause on an AI nonconformity usually lands in governance rather than engineering. A missing impact assessment for a new feature is rarely because the method was too hard. It is because nothing in the release process asks whether an assessment is needed, or because "new feature" and "new AI system" are not the same event in anyone's process.
That changes what you write. "Engineer to complete the assessment" is a correction logged as a corrective action, and the finding returns next quarter under another system name. "Add an AI-system trigger to change approval, with the AIMS owner as a required reviewer" is corrective action. Keep the chain documented — nonconformity, correction, cause analysis, action, and the later check that it worked.
Management Review as the Other Half of Clause 9
Management review is the other half of ISO 42001 clause 9, and where the audit stops being an artefact and starts changing decisions. Leadership examines the AIMS at planned intervals, weighs what it is given, and decides what to change: resources, objectives, scope, or the system itself.
What it should receive from the audit is not the report. It is the parts requiring a decision above the AIMS owner's authority — nonconformities that cannot close without budget, findings that question the scope, trends across audits, and the systems the audit could not get adequate evidence for. A reviewer handed a 40-page audit file approves it and changes nothing.
The management review AI governance needs also looks forward: risk and impact changes since the last review, vendor model changes landing next quarter, AI features on the roadmap, and whether the scope statement will still be accurate in 6 months. ISO publishes the standard that describes the cycle; what makes it work is leadership treating the review as a decision meeting with minuted outcomes rather than a status read-out.
ISO 42001 Internal Audit Questions Teams Ask
These are the ISO 42001 internal audit questions teams ask most often when planning a first audit cycle against a live AI estate.
At planned intervals, with the frequency set by your audit programme rather than by a number in the standard. Most organizations run a full pass across the AIMS scope annually to feed management review and the certification cycle, then audit higher-consequence systems more often. Base the frequency on the importance of the processes and the results of previous audits.
Not for the parts they are accountable for. Independence means auditors do not audit their own work, and the AIMS owner auditing the AIMS is the arrangement a Stage 2 auditor queries first. Small organizations resolve it by splitting the estate so nobody audits a system they built, pairing an independent lead auditor with a technically competent reader, or engaging an external auditor to run the internal audit.
The procedure is nearly identical — programme, plan, fieldwork, findings, report, corrective action. The difference is that the audited environment moves, so evidence has to be dated to a model and dataset version and the audit has to test the processes that absorb change rather than only the artefacts change leaves behind. The subject differs too: impact on individuals and society, data provenance and oversight have no real ISMS equivalent.
An incomplete AI system inventory, followed by a Statement of Applicability that has not moved while the estate has. Both share a root cause: AI systems arrive through routes the governance process does not watch — an embedded vendor feature, a model added to an existing product, a team's own tooling. Testing the inventory from outside rather than reviewing it from inside is the highest-yield test available.
Every cycle has to cover the AIMS scope, but a single audit does not have to cover everything at once. A programme can split coverage across several audits in a period, provided the whole scope is covered within the cycle and the split is driven by process importance and prior results. A certification body looks for a programme that demonstrably covers the system over time.
Run the AIMS audit cycle on evidence that stays current
Book a demo and we will show how Konfirmity keeps the AI inventory, impact assessments and model change history audit-ready between internal audits.
Book a demo
Audit the System, Not the Snapshot
The internal audit that earns its cost tests whether the AIMS still works after the estate changes, because that is the only condition it ever operates under. Conformity against a checklist is a day's work and tells leadership very little. Effectiveness against a moving estate takes longer, produces uncomfortable findings, and is what a Stage 2 auditor will test anyway.
Start with the 4 questions. Test the inventory from outside it, date the Statement of Applicability against the estate, look for one real oversight override, and ask what happened the last time a vendor moved a model underneath you. If any of those produce a shrug, you have your first finding — and you have it before an external auditor does.
Then make the programme compound: carry this year's weak areas into next year's plan, send management review the decisions rather than the document, and write corrective actions that change a process. If the underlying artefacts are still being assembled, the ISO 42001 checklist is the place to start, so the audit tests something real instead of reviewing intentions.







