Nothing to hand? Load the — a quarter of payment releases where the same person prepared and approved several runs, two invoices were paid twice, and three amounts sit just under the authorisation limit — or the , a monthly reconciliation control with complete evidence and nothing to report. Both replay a saved run for free.
Sizing and selecting a sample is arithmetic, and it is where workpapers actually fail
The sample size comes from the control's frequency and its risk of failure, adjusted for a prior-year deficiency, for external auditor reliance, and for whether it is a key control — and here every adjustment is its own line with its own delta and its own stated reason, so a reviewer can argue with one without re-deriving the rest. An automated control gets the one-instance override instead of the frequency base, never both: applying both is the double count that quietly doubles a sample size and survives review because nobody re-adds the column.
A random selection nobody can reproduce is not evidence
Selection runs in three disjoint passes: everything at or above the high-value threshold is tested at 100% and removed from the pool; then risk-targeted items are drawn against the attributes the browser found; then the statistical sample is taken from what is left, systematically for a periodic control and randomly for a transactional one. Because the pool shrinks as each pass claims items, no item is ever counted in two strata — and the page asserts that rather than assuming it, since an item counted twice overstates both count and value coverage. The whole draw is seeded, so re-entering the same population, options and seed reproduces the identical selection: that is what makes a random sample auditable.
The metered pass is judgement, and it is allowed to disagree
Twelve risk attributes are matched against the population — the same person preparing and approving, blank approvals, duplicate references, exact offsetting pairs, amounts sitting just under an authorisation limit, weekend postings, cut-off items, sensitive accounts — and each test reads only the fields it is about, because a memo line containing the word “approved” must never satisfy an approver test. Enter your exceptions and the browser computes the deviation rate, the tolerable rate, a rule-of-three upper bound where there were none, and the extrapolated error against your materiality. What it cannot do is decide whether one deviation was isolated, whether a compensating control catches it, or whether the failure was of the control or only of its documentation — so the model is asked for exactly that, plus test steps written for your control as you described it. It must return a verdict on every attribute, and calling one a data artefact is often the right answer: a duplicate reference is usually a reversal. Workpapers are saved to your account, so next quarter's re-test is diffed against this conclusion.
This app assists with a SOX testing workflow. It is not audit or legal advice. Every workpaper, sample and assessment it produces must be reviewed by a qualified financial professional before it is relied upon or included in audit documentation. Sample sizes and the significant-deficiency threshold used here are this app's stated conventions, not requirements of any standard. A derived work of @anthropics/sox-testing, whose control matrix, sample-size table by control frequency, selection methods, workpaper template and deficiency classification framework are what this app measures against.