Internal Controls Testing: A Practical Playbook for 2026

Issabelle Fahey

Issabelle Fahey

Head of Growth
27 July 2026

The first time internal controls testing lands on your desk, it usually isn't because someone woke up excited about audit hygiene. It's because a fractional CFO said “SOX-ready,” a bank asked for evidence you don't have in a neat folder, or an acquirer started treating your close process like a crime scene. Fun times, right?

At that point, the game changes. You're no longer just trying to keep the books tidy. You're building a system that can survive scrutiny, keep the team from firefighting every month-end, and stop the close from turning into mortgaging your office ping-pong table for one more workaround.

Internal controls testing is the discipline that proves your controls exist, work, and leave evidence behind. Done well, it makes fundraising easier, reduces ugly surprises, and stops operations from running on heroic spreadsheets and vibes. Done badly, it becomes a side quest that eats the finance team alive.

When Internal Controls Testing Becomes Your Problem

The moment it becomes your problem is usually the moment “we'll clean that up later” stops being believable. A founder who could once hand-wave small process gaps suddenly gets asked for control evidence, and now every shortcut has a paper trail or a missing one. That's when internal controls testing stops sounding like audit jargon and starts sounding like survival.

Here's the plain-English version. You pick the controls that matter, check whether they're designed sensibly, then test whether they ran the way people claim they did. That matters because internal control isn't a one-time policy doc, it's the way the business keeps transactions accurate, authorized, and documented while people are busy doing real work. COSO-based guidance treats control as an ongoing process, not a binder on a shelf, and practical control-testing guidance says the sequence is understanding, design, implementation, and evidence, not guesswork (Pathlock's practical guide, GAO-oriented guidance summarized by Wolters Kluwer).

Why it shows up all at once

Growing companies hit this wall for a few predictable reasons. Finance is closing faster, systems are multiplying, and more people can touch the numbers without anyone meaning to create risk. A control that worked when one person owned AP and the bank rec becomes flimsy the minute there's a second approver, a new ERP, or a remote ops team.

That's also where the payoff starts. Good testing doesn't just keep auditors calm. It reduces repetitive reconciliation drama, makes closes cleaner, and gives investors less room to panic over process gaps. For public companies, the resource burden is real. KPMG's 2023 SOX report says the average SOX program budget was $1.6 million, the average program took 11,800 hours, and the average compliance cost was $3,200 per control (KPMG SOX report). That's not a spreadsheet hobby. That's a real operating cost.

Practical rule: if your controls only exist in someone's head, you don't have controls. You have optimism with a login.

The upside is simple. When you test controls properly, you catch issues before they become investor questions, board headaches, or restatement nightmares. You also build a cleaner operating rhythm, which holds significant value that often goes unrecognized.

Scoping and Risk Assessment Without the Theater

Skip the committee cosplay. Scoping starts with one question, which processes can move the financials, the compliance posture, or the fraud risk needle? If a control does not affect reporting, asset protection, or access to money and systems, it probably does not belong at the top of the list. That is not laziness. That is triage.

A useful scope starts with the processes that can hurt you fastest, revenue recognition, AP, payroll, journal entries, access management, and any interface that moves data between systems. Then rank controls by impact, likelihood, complexity, and fraud risk. That is the founder version of risk assessment, and it is far more useful than a meeting where everybody says “critical” about everything.

What gets tested first

The controls that deserve attention first are the ones where failure creates real mess. A manual journal entry review that can be bypassed gets priority. A monthly reconciliation with a named reviewer and an evidence trail also matters. A low-risk checklist that only repeats a system-generated report usually is not where you want to spend your best people's time.

You decide which controls are key and which ones are supporting cast. Key controls are the ones you need to rely on. Compensating controls matter too, but they should not become an excuse to test every dusty checklist in the building. Diligent's guidance points to a risk-based testing cadence, with high-risk controls often tested monthly or quarterly, medium-risk controls semi-annually, and low-risk controls annually (Diligent controls testing guidance). That is the right mindset, risk drives cadence, not calendar theater.

The Australian Tax Office makes the same practical point. If a key tax control has never been tested, you need to map its frequency and the assumed population first because standards do not hand you fixed sample sizes (ATO control testing guide). Translation, do not make up a sample because someone is in a hurry.

Use this as your one-page scope:

Scope Question Decision Rule Founder-Friendly Answer
Does the process hit financial reporting? Include if yes Keep it in scope
Can one person initiate and approve? Include if yes Test it hard
Is the control manual or system-driven? Manual needs more scrutiny Give it more evidence
Does failure create fraud or restatement risk? Prioritize if yes Move it up the list
Is there a compensating control? Use it only if reliable Do not overtrust it

A diagram illustrating the five evidence methods for internal controls, including inquiry, observation, inspection, re-performance, and CAATS.

Your scope memo should do two jobs. It should make your life easier, and it should make it hard for an external auditor or investor to redraw the map later. If you can explain why a control was in or out in two minutes, you probably scoped it well.

This accounts payable process guide is a useful companion if AP is one of your messiest areas, because that is where a lot of scope creep starts.

The Five Evidence Methods and When Each One Actually Works

A startup can get through one audit with sloppy evidence and still survive. The second audit is where the shortcuts start breaking. Inquiry, observation, inspection, re-performance, and CAATs each answer a different question, and if you use them interchangeably, you end up with pretty binders and weak proof.

The hierarchy of proof

Inquiry is where you start, not where you finish. It helps you map the process, identify the owner, and find the records that should exist. It does not prove the control ran. If someone says “yes, we do that every month,” you have a lead, not evidence.

Observation is stronger, but only for controls you can watch. Seeing a reviewer sign off on a close checklist shows the step happened in front of you. It still does not prove the rest of the period looked the same. Inspection is better because you are reviewing records, approvals, logs, or reconciliations. Even then, a clean PDF can be staged, so do not treat inspection like a guarantee.

Re-performance is the one that matters most when you want real proof. The ATO testing controls guide says it provides the strongest evidence for operating effectiveness, and that matches how these tests work in practice (ATO testing controls guide). If you can run the control again and get the same result, you have something real. That is why re-performance works so well for reconciliations, access reviews, and close controls.

CAATs, or computer-assisted audit techniques, are the right answer when the process runs through a system and leaves a data trail. MetricStream points out that CAATs can test the full population instead of a sample (MetricStream control testing overview). If your ERP or workflow tool already has the data, stop pretending a handful of manual samples tells the whole story.

Inquiry tells you where to look. Re-performance tells you whether the control deserves to stay in place.

Digital evidence changes the game

Older testing habits get awkward fast once the evidence lives in systems instead of filing cabinets. A control might leave behind an ERP log, a workflow timestamp, a ticket in an access review tool, or an approval sitting inside finance software. The method does not change, but the evidence source does.

The strongest test is usually layered. Use inquiry to understand the process, inspection to collect the record, and re-performance or CAATs to prove the control held up. If a control is automated, system-generated evidence should carry more weight than someone's memory and a screenshot from late Friday.

Use the evidence that matches the control. If the control lives in a system, test the system evidence. If it is manual, test the human step and the output. If volume is high, use CAATs and stop treating tiny samples like they settle anything.

Sampling Strategies That Match Your Control Frequency

Sample size isn't a vibe. It's a function of population, frequency, and risk. Wolters Kluwer's guidance says sample size should be based on the control's population and frequency, which is exactly why monthly, quarterly, and annual controls should not be treated the same (Wolters Kluwer internal control testing guidance).

The practical move is to match how often the control runs with how much evidence you can reasonably review. A monthly control gives you more opportunities to test, but each period often has a smaller population. An annual control usually has a larger or more consequential population event, which changes the logic. Small populations often need broader coverage because there isn't enough volume to hide behind fancy sampling language.

A simple reference table

Control Frequency Typical Population per Period Recommended Sample Size Notes
Monthly Low to moderate 1 to 3 items, or targeted full coverage if volume is tiny Use when the control is high risk or the population is small
Quarterly Moderate 2 to 5 items Good for recurring reviews, approvals, and reconciliations
Annual Low count, high impact Full coverage if feasible, otherwise test the key instance and all exceptions Best for infrequent controls with limited occurrences

Don't treat that table as a magic formula. Treat it as a planning rule for SMBs and startups that don't have the luxury of building a mini internal audit shop on day one.

Judgmental or statistical

For scrappy teams, judgmental sampling is usually the right call. You choose the items that look most risky, most representative, or most likely to expose a design flaw. Statistical sampling makes sense when you have stable, high-volume populations and need a more formal conclusion. Most startup finance teams don't live there yet.

If the population is tiny, test more of it. If the population is messy, target the unusual items. If the control is automated and the system can export a complete population, skip the guessing game and use the full dataset where you can.

The cleanest defense to an auditor is short. Say what the population was, how often the control runs, why the sample was enough, and what you did with exceptions. That's it. If you start talking like a textbook, you probably don't trust your own sampling decision.

Running Test of Design and Test of Operating Effectiveness

A control that looks good in a policy can still fall apart in real use. Test of design asks whether the control, on paper, can prevent or detect the risk. Test of operating effectiveness asks whether it ran the way it was supposed to, with the evidence to prove it.

Start with the walkthrough, then test design, then test operating effectiveness, as outlined in the Wolters Kluwer internal control testing guidance. That order matters because founders often approve a control framework that sounds clean but breaks the moment the team gets busy, the close gets rushed, or someone takes a shortcut and never writes it down.

A close-process example

Take a month-end close control. The controller reviews the revenue reconciliation, checks the exception report, and signs off before the books close. On paper, that sounds fine. For TOD, ask whether the reviewer has the right data, whether the review happens before the close lock, and whether the control can catch a misstatement.

For TOE, test a real period and use real evidence. Look for the sign-off, the report, the exception handling, and proof that the reviewer investigated a problem. A walkthrough should still include inquiry, observation, and inspection. If the control matters, re-perform it with the underlying data so you are not relying on a screenshot that looks tidy but proves very little.

The same guidance also points to understanding the control, assessing whether the design works, checking that it was implemented properly, and verifying that transactions are documented in a way you can trace later (Wolters Kluwer internal control testing guidance). That sequence is the whole job. Skip a step and you usually end up defending a test that does not hold up under review.

Practical rule: if the control cannot be re-performed or traced to system evidence, expect the evidence to be challenged later.

How to write the test result

Write the result like someone who has had to defend it in a board meeting. Say what you tested, what period you tested, what evidence you inspected, and whether the result matched the expected control behavior. If it failed, document whether the issue is a deficiency, a significant deficiency, or a material weakness under your company's reporting framework. Those labels drive who gets notified and how quickly the fix has to move.

Keep the memo short and factual. A strong test memo gives the reader enough detail to follow the logic without having to guess what happened. If the reviewer cannot recreate your reasoning from the document, the test was not documented well enough.

Use the same standard when you tie the test back to the broader finance process. Clear control write-ups make it easier to connect testing results to financial reporting best practices without turning the file into a pile of loose notes.

Documenting Findings, Remediation, and Retesting

A finding that doesn't drive action is just expensive journaling. The useful version names the issue, the root cause, the owner, the fix, and the retest trigger. If you leave any of those out, the finding tends to live forever in a folder nobody opens until the next audit panic.

The easiest way to ruin remediation is to treat it like a ticket instead of a management problem. The controller finds the gap, engineering says the fix will slow a release, and suddenly everyone is debating priorities instead of closing a control weakness. That's normal. It's also exactly why the documentation has to be sharp enough to survive pushback.

What good remediation documentation looks like

Write the exception in plain English. Then state why it happened, not just what happened. A weak segregation-of-duties control, an approval flow that skipped a required step, or a reconciliation performed too late all need different fixes. The action plan should have a named owner and a due date, not a vague “in progress” that somehow lasts until summer.

The Blue Sage internal control testing guidelines say testing should be performed in at least two cycles to allow remediation, and controls should be retested after remediation is complete and after enough of the population has accumulated to make the retest meaningful (Blue Sage guidelines). That's the right call. Retesting too early just proves the fix exists in theory. It doesn't prove the control works in real life.

The retest rhythm

Retest only after the fix is live and enough transactions have flowed through the control to make the retest meaningful. If you retest the day after the patch lands, you're mostly testing enthusiasm. You want evidence, not a victory lap.

This financial reporting guide is a useful companion when the finding touches close quality, because a lot of control failures show up there first.

If the remediation plan doesn't survive a quarter, it wasn't a plan. It was a hope with an owner field.

The best teams treat retesting as part of remediation, not as optional housekeeping. That shift keeps the control environment honest and stops old problems from reappearing under a fresh label.

SOX, GAAP, and the Founder's 30-Day Action Plan

SOX, GAAP, and controls testing get mashed together so often that founders start thinking they're one giant compliance soup. They're not. GAAP drives how financial reporting should be prepared, while SOX raises the bar for public-company internal control over financial reporting, especially once Section 404 becomes relevant. The smart move is to build a control system early, before a public-company calendar forces you to.

The practical trigger isn't just an IPO. It's also when your growth makes sloppy controls expensive. If you're two years out from a likely public listing, or if investors are already asking for audit-ready evidence, now's the time to stop pretending spreadsheet heroics count as governance. The internal control framework needs to be real before the pressure becomes public.

A 30-day action plan that won't waste your time

Week 1, inventory the mess. List the financial processes, owners, and current control evidence. Identify where approvals, reconciliations, access, and journal entries live. If the accounting stack is changing, check the relevant standards update path through accounting standards updates so the control design isn't built on stale assumptions.

Week 2, rank the risks. Put the highest-risk controls at the top, the ones tied to reporting, cash, access, or close. Decide what gets tested monthly, quarterly, or annually, and be ruthless about low-value controls. If a control doesn't matter, don't waste testing time proving it's boring.

Week 3, run TOD and TOE on the key controls. Do the walkthroughs, collect the evidence, and re-perform the toughest ones. Use the five evidence methods like tools, not rituals. If a control can be supported by system logs or CAATs, use them.

Week 4, lock remediation and retest. Write findings, assign owners, and schedule the retest only after the fix has had enough time to operate. That's how you avoid fooling yourself with a fresh patch and no proof.

For a strong external checklist perspective, flawless compliance audit tips is a decent reference point, especially if you're trying to pressure-test your own prep before someone else does.

Build or buy

If you're early, outsource the heavy lifting and keep the judgment in-house. A remote senior accountant or internal audit contractor can buy you speed without forcing you into a bloated team too soon. When the control environment starts scaling across entities, systems, or jurisdictions, then you start thinking about a fuller SOX team.

My blunt take? Hire for judgment first, muscle second. A sloppy control program with more headcount is still sloppy. A lean program with clear owners and good evidence wins every time.


If your finance team needs help building a controls testing program, tightening close discipline, or hiring people who can execute the work rather than just discuss it, HireAccountants can help. They place pre-vetted accounting and finance talent fast, which is exactly what you want when controls are getting real and your team needs sharp hands, not extra meetings.

Ready to streamline your accounting?

Let's simplify your finances today!