Agent Operating Guide
/sox-testing
Plan and drive one control area's SOX 404 testing for a period, from control matrix through sample selection, annotated workpaper, and deficiency evaluation, with a fresh-context grade at the end.
The front door of the SOX fieldwork engine. Given a control area and a period, it plans and drives the whole test: it builds the control matrix, sizes the sample, validates the population, selects and draws the sample, assembles the workpaper, classifies any exceptions, and finishes with an independent grade. Population (IPE) validation and cross-evidence consistency checks are built into the fieldwork rather than bolted on. It delegates rather than working inline: deterministic steps go to /sox-python, and evidence annotation goes to the /sox-annotate-xlsx, /sox-from-video, /sox-from-web, and /sox-from-folder helpers.
Inside AssureSwarm you rarely call this directly: /sox-test is the bridge
that pulls a workflow step’s context and gathered evidence, runs this engine, and lands the
result back in Canvas for review. Run /sox-testing on its own when you’re working outside
a tenant or driving the fieldwork by hand.
When to use
Section titled “When to use”When you’re planning quarterly or annual SOX 404 testing, pulling a sample for a control (revenue, procure-to-pay, ITGC, financial close), validating a population before sampling, building a testing workpaper, running a period’s tick-and-tie end to end, or evaluating and classifying a control deficiency.
Inputs
Section titled “Inputs”| Flag | Required | Notes |
|---|---|---|
<control-area> |
Yes | One of revenue-recognition, procure-to-pay (p2p), payroll, financial-close, treasury, fixed-assets, inventory, itgc, entity-level, journal-entries, or a specific control ID/name. |
<period> |
No | The testing period (e.g. 2024-Q4, 2024, 2024-H2); the skill asks if absent. |
--samples <file> |
No | A tester-supplied sample list (csv/xlsx/json) to use instead of drawing one; its provenance is still recorded in a procedure tab. |
--prior <workpaper.xlsx> |
No | Last period’s workpaper, read-only, to seed this period’s plan; deltas are confirmed with you before the plan is written. |
--evidence <dir> |
No | A folder of mixed evidence files, routed to /sox-from-folder once the plan and samples are locked. |
Example
Section titled “Example”/sox-testing procure-to-pay 2024-Q4 presents the P2P control matrix, sizes the sample
against the methodology reference, validates the population, and draws the sample through
/sox-python with a recorded seed. It then locks each attribute’s expected value, extracts
and annotates the evidence, auto-judges every field, and grades the finished workpaper:
handing you a Summary matrix with a written reasoning cell on every row.
How it works
Section titled “How it works”The engine builds a test plan, or, with --prior, seeds one from last period’s workpaper
and confirms the deltas with you. It identifies the key controls, sizes the sample against
the audit-support methodology, validates that the population is complete, and selects the
sample, delegating the actual draw to /sox-python so the seed is recorded and
reproducible. It locks each attribute’s expected value and commits the exception policy
before any results exist, then routes evidence to the right helper, whose leaf agents, a
boxer that locates each field, a context reader that reads its value, and a reviewer that
double-checks box position, keep the image bytes out of the engine’s own context. It
auto-judges each field pass/fail behind four gates or routes it to needs_human, runs a
cross-evidence pass for duplicated evidence and segregation-of-duties conflicts, and finally
dispatches the sox-workpaper-grader leaf agent to grade the finished workpaper in a fresh
context window: that isolation from the engine’s reasoning is what makes the verdict
meaningful. On a passing grade it recommends /sox-replay-build.
Good to know
Section titled “Good to know”- Every deterministic step routes through
/sox-python. Sample draws, three-way matches, recomputations, and threshold checks all run as code so their source, output, and a script SHA-256 land in a reproducible procedure tab: a reviewer can re-run them. - Auto-judged verdicts clear four gates or defer. A pass/fail is written automatically
only when the observed value unambiguously matches the expectation, the read is confident,
the box position is confirmed, and no disqualifying context signal (disabled, unsaved,
error, out-of-force status) fires. Anything else is surfaced to you as
needs_humanbefore grading. - The exception policy is committed at plan time. Choosing expand-vs-stop after seeing the first exception is outcome-shopping, and the grader checks for it.
- Zero-occurrence populations are handled. A control with no opportunity to operate is marked “Not applicable: no occurrences,” with the validation kept as proof.
- Reasoning is mandatory. A Summary row without a reasoning cell is a deficient workpaper; the engine never leaves one blank.
- The canonical writer stays legacy; the grader reads both shapes.
/sox-testingkeeps writing ones<N>_<TestSlug>detail tab per (sample, test) pair. The grader also accepts a record-centric workpaper, where Summary comes first, there is exactly one tab per sampled record, and every attribute for that record sits on that record’s tab in a six-column result table (Attribute,Result,Observation,Procedures Performed,Conclusion,Evidence). It follows the Summary hyperlinks into that table instead of looking for detail tabs, and it counts the annotation shapes each form uses: the injectedsox-*rectangles on the legacy side, the independently movabletickmark-label-*callouts on the record side. - Record-centric links are validated to the exact row. The
Sample #cell and every result cell in a Summary row must point at the same record worksheet, and each result link must land on that test’s exact attribute row. A link to the right sheet but the wrong row is a failure, not a rounding error. - A mixed workbook fails closed. The shape is detected once, from the Summary’s own
structure: the header row is the single row carrying both
Sample #andReasoning(row 1 in a legacy workpaper, row 6 in a record-centric one that opens with a preamble block), and two candidate rows or none is an error rather than a guess. A workbook that mixes both forms, or matches neither, is never graded as if it were one of them: the structural criteria fail and the detector’s own errors are quoted as the evidence.
Related
Section titled “Related”- /sox-test: the AssureSwarm bridge that runs this engine on a workflow step.
- /sox-python: the deterministic runner every draw and tie-out delegates to.
- /sox-from-folder: where a mixed
--evidencedirectory is routed. - /sox-replay-build: package a passing run so it repeats next period.
Not audit or legal advice. Workpapers and assessments produced by these skills require review by qualified financial professionals before being relied on for SOX 404 compliance.