All articles
8 October 2026

More Assurance for Less Effort: How AI Transforms Control Testing

AI control testing automates the management testing your teams already do. It cuts the hours spent chasing evidence and raises assurance by testing every item, not just a sample.

AI audit pipeline with human review

Why it matters: less effort, more assurance

Every year, companies spend a large share of their compliance budget proving that their controls work. Most of that effort goes into chasing, collecting and reading evidence rather than judging it. AI control testing changes the equation: it automates the management testing that compliance teams and external testers already perform, so the same team can test more, finish sooner, and stand behind every conclusion with stronger evidence. It pays off in two ways.

It makes testing far more efficient. Most testing hours go to requesting, uploading and reading evidence, not to judgement. AI takes over those steps, so a test cycle that used to drag on for weeks of reminders can finish in days, and testers spend their time on the exceptions that actually need a human.

It gives you more assurance, not less. Manual testing checks a sample and hopes it is representative. AI can test every item in the population, apply exactly the same attributes every time, and cite the evidence behind each result. You find the failures a sample of 25 would have missed, and an auditor can trace every conclusion back to its source.

Same process, far fewer hours, and every item testedTodayScopeSample 25RequestChase every ownerUploadScreenshots by handReviewRead every itemConcludeSign offCoverage: a sample · Effort: weeks, mostly spent chasing evidenceWith AIScopeAttributesAI pulls and testsEvery item, with citationsReviewExceptions onlyConcludeSign offTime backfor judgementCoverage: every item · Effort: days, focused on the exceptionsMore efficiencyHours shift from chasing evidence to judgementMore assuranceEvery item tested, every result traceable

Adoption is still early. In a 2026 Gartner poll of 743 audit professionals, only 30% used AI for audit testing (Gartner via IT-Online). This article covers how management testing works today, how AI automates it, the tools available and the architectures that hold up in front of an auditor.

How management testing works today

Management testing is the periodic check, by a compliance team or an external tester, that key controls actually operated. Most programmes follow the same five steps:

  1. Scope is defined and announced. The testing team selects controls, period and samples, and notifies control owners.
  2. Evidence is requested. Control owners receive a request for each sampled instance: tickets, approvals, access lists, screenshots.
  3. Evidence is uploaded. Owners upload it to a repository, ideally the organisation's GRC system, where it is linked to the control and the test.
  4. Evidence is reviewed. Testers check each item against the test attributes to decide whether the control was performed correctly.
  5. The control is concluded effective or ineffective. Deficiencies become issues with owners and remediation dates.

In Swedish companies, step 3 is typically manual. Control owners export reports and take screenshots and upload them by hand, usually the day after the deadline and named something like screenshot_final_FINAL2.png. Automatic extraction from source systems is the stated goal almost everywhere, but rarely the reality. Most hours therefore go to requesting, uploading and reading evidence rather than to judgement.

AI testing automates the same process

AI control testing keeps the five steps and changes who, or what, performs them. Scoping and the conclusion stay with people; the evidence-heavy middle is where automation pays.

AI takes over the evidence-heavy middle; scoping and sign-off stay humanTodayWith AI1 ScopeTester scopesControls, period and samplesannounced to control ownersTester scopesAI drafts the test attributes;full population, no sample2 Request evidenceTester requestsOne request per sampled itemReminders by email, repeatedlyRequests only the gapsEvidence that cannot be pulledReminders sent automatically3 Upload evidenceOwner uploadsScreenshots and exportsuploaded by hand to GRCConnectors extractFrom ITSM, IAM and ERP;manual upload where no API4 Review evidenceTester reviewsReads every item againstthe test attributesAI tests every itemResult, confidence, citation;tester checks the fails only5 ConcludeTester concludesEffective or ineffective;issues raised for gapsTester concludesNamed person signs off;the AI is never the authorAutomated by AIStays with a person

Steps 2 to 4 are where the hours go, and they are exactly the steps AI takes over.

Automation does not have to start with extraction. The review step is usually the largest, and AI can perform it on evidence that owners have uploaded by hand. That makes manual-upload environments, common in Sweden, a realistic starting point rather than a blocker.

Each automated step must leave a trace: prompts, model version and evidence hashes, so an auditor can re-perform the test.

What AI actually tests: examples

The best candidates are controls with clear yes-or-no attributes and plenty of instances, which is exactly where human testers get bored and start to skim.

Control What the AI checks Typical evidence Fit today
Change and deployment Approved before deployment; approver is not the developer; testing documented; emergency changes approved afterwards ITSM tickets, CI/CD pipeline logs High
User access review Review completed on time; every account marked for removal actually removed IAM exports, review sign-off High
Leavers Access removed within the agreed deadline after termination HR leaver list against AD and ERP users High
Privileged access Admin accounts approved, justified and reviewed AD groups, PAM logs High
Backups Jobs completed; restore tests performed Backup logs, restore test records High
Manual journal entries Entries above threshold approved by someone other than the preparer ERP ledger, e.g. FGLEDG in M3 High
Vendor assurance SOC report obtained, in period, exceptions assessed Uploaded SOC reports Medium
Management review Review performed with sufficient precision Meeting minutes, review packs Low: AI checks completeness, not judgement

Take change management as the worked example. Today a tester samples, say, 25 changes and asks owners for the tickets. With AI, every change in the period is pulled from the ITSM tool and checked against four attributes, and the tester only looks at the handful that fail. The AI does not get bored on ticket number 400, and it does not need coffee.

One change, four attributes, one exception for a human to look atCHG0041873 · Deploy payment service v2.3Illustrative exampleAttributeResultEvidence citedApproved before deploymentPassApproval 09:12, deployment 14:30Approver is not the developerPassApprover K. Lind, developer A. BergTesting documentedFailNo test record linked to the ticketEmergency change approved afterwardsN/AStandard change, not an emergencyOutcome: exception routed to a tester, who confirms or overrides with a reason

Each result cites the evidence it rests on, so the tester can verify the one failure in seconds instead of re-reading the whole ticket.

Available tools

Most large GRC platforms now automate part of the review step, but several headline features are announced rather than available. Status is as of October 2026; performance figures are vendor claims.

Tool What it automates Status
LogicGate Spark AI Reviews evidence; returns pass, fail or incomplete with a rationale Available
Optro (formerly AuditBoard) Autonomous Testing Attribute tests, access reviews and reconciliations from raw evidence Available
Hyperproof AI Evidence collection and validation Available
Vanta AI Agent Evidence evaluation for SOC 2 and ISO 27001 Available
Archer Evolv Scoped AI agents with a named human supervisor Available
ServiceNow IRM with ComplianceCow Native test workflows; partner app turns control text into executable tests Available
Workiva Automated Testing Evidence, attribute and testing agents Announced, no GA date

Spotlight: Optro Autonomous Testing

If one product shows where control testing is heading, it is Optro, the platform formerly known as AuditBoard, which renamed itself in March 2026 (Internal Audit Guide). In May 2026 it acquired Midship and turned it into Autonomous Testing: agents that take raw evidence and perform attribute tests, access reviews and reconciliations end to end.

Three things make it worth a closer look:

  • It tests, not just assists. The agent works through each attribute and documents its result, so the tester reviews outcomes instead of reading every ticket.
  • It covers the whole cycle. An earlier Audit Agent handles risk-based sampling and evidence annotation, and Document Intelligence turns walkthrough notes into control narratives.
  • It is open to other agents. An MCP server, launched in April 2026, lets external AI agents read and write GRC data under the platform's permissions.

The caveats are the usual ones. Midship's claim of automating up to 87% of SOX programme management is the vendor's own figure, and the product was built for US SOX teams. Swedish buyers should ask about EU data residency and about how well the agents handle evidence in Swedish.

Building your own is also an option: an LLM, plus a GRC system that exposes an API or MCP server, can automate review for your highest-volume controls. It is cheaper per control but puts the validation burden on you.

Possible architectures

Whatever the product, a sound AI testing architecture has the same five layers, with governance running alongside every layer that touches evidence or results.

Every AI test result passes human review before it reaches the GRC recordresults, confidence, evidence citationsapproved conclusionsSource systemsITSMIAM / ADERPCloud configEvidence layerExtracted from source systems or uploaded by owners; hashed and timestampedTest engineDeterministic rulesFull-population checks onstructured dataLLM agent testerReads tickets, approvals andscreenshots; cites evidenceHuman reviewTester triages fails and low confidence; reviewer signs offGRC system of recordControls, test results, issues: ServiceNow IRM, Archer, OptroGovernanceAgent inventory and ownersLeast-privilege accessVersioned prompts, modelsand rulesEvery action loggede.g. ServiceNow AIControl Tower, OneTrust

The key design choice is that the AI proposes and a person concludes. Where evidence is still uploaded by hand, the GRC repository simply acts as the evidence layer; the rest of the architecture is unchanged.

Four patterns implement these layers, and most programmes end up combining two of them.

Pattern How it works Strengths Weaknesses Best fit
Embedded platform agent The GRC vendor's own agents test inside the platform (Optro, LogicGate, Archer, Workiva) Fastest start; audit trail built in Limited to vendor connectors; model is a black box; lock-in Teams standardised on one platform
Deterministic rules generated by AI A model writes the test rule or code once; the rule then runs without AI (ComplianceCow, SAP Joule) Repeatable and re-performable; cheap at scale Structured data only; the generated rule still needs review ITGC, configuration, ERP transactions
External agent via MCP or API Your own agent reads evidence and writes results into the GRC platform through MCP servers or APIs (ServiceNow Action Fabric, Optro MCP, OpenPages MCP) Free choice of model; covers unusual controls You own validation, security and change control Teams with engineering capacity
Multi-agent pipeline Separate agents collect evidence, test attributes and challenge results, under an orchestrator Mirrors segregation of duties; a checker agent catches errors Complex, costlier, harder to explain to auditors Large, high-volume programmes

On ServiceNow, a practical combination is native indicators and control tests for structured checks, AI Agent Studio or a partner app for evidence reading, and AI Control Tower for agent inventory and access.

Making the results auditor-ready

No audit standard setter has written AI-specific testing rules, so existing, technology-neutral rules apply in full. The PCAOB has no generative AI standard, but its amendments on technology-assisted analysis apply to fiscal years beginning on or after 15 December 2025 (Fieldguide). The IIA Global Internal Audit Standards likewise require due care and sufficient evidence whatever the tool.

In practice, an external auditor relying on AI-tested controls will ask four things:

  • Is the data complete and accurate? Show where the population came from and how you reconciled it.
  • Is the test re-performable? Keep prompts, model version, rule code and evidence hashes with each run.
  • Who concluded? A named human signs off. ServiceNow's AI Response Assist is a useful pattern: the AI suggests, the user applies, and the AI is never recorded as author.
  • Is the tester itself controlled? A test agent is part of the control environment. It needs change management, its own test suite, and segregation between whoever builds it and whoever approves its output.

EU rules add a layer. A testing agent is an ICT service, often from a third party, so it belongs in the DORA register of information and NIS2 supply-chain assessments, with least-privilege access and every action logged.

Getting started

Start where management testing hurts most, and widen the scope once your auditor accepts the method.

  1. Automate review of uploaded evidence first. It needs no integrations and attacks the largest share of testing hours.
  2. Pick five to ten pilot controls with clear attributes, such as access reviews, terminations and change approvals.
  3. Run in parallel for one cycle. Compare AI and manual results and measure false passes and false fails.
  4. Agree the evidence pack with your external auditor before go-live.
  5. Add automatic extraction control by control, as source-system connectors become available.

Avoid two common traps, both more tempting than a Friday fika: treating a vendor's accuracy figure as your own, and automating judgement-heavy controls too early.

Closing thoughts

AI control testing is management testing with the evidence work automated. Testers move from requesting and reading evidence to designing attributes, reviewing exceptions and concluding. The technology is ready for structured, high-volume controls today; what most programmes still need to build is the method that lets an auditor see exactly what the AI did and why. The reward: testers get their weeks back, and control owners might finally get through July without a single evidence request.