Skip to content

Agent Operating Guide

/sox-from-video

Turn a recorded auditor walkthrough and its transcript into annotated SOX workpaper tabs, with matching video frames boxed by movable red Excel shapes over each tested field.

Takes a recorded auditor walkthrough, a Zoom, Teams, or Loom mp4, plus a transcript and produces the same SOX workpaper format the other evidence skills do: a Summary tab over per-sample-per-test detail tabs, each carrying the red-bordered narrative on top and annotated video frames below, with red-rectangle Excel shapes over each tested attribute (movable shapes, never burned-in pixels). The skill is self-contained: frame extraction is vendored, and vision is the model itself, no OCR.

It never reads the transcript or views a frame itself. A sox-walkthrough-parser turns the transcript into a list of review-moment timestamps, a vendored extractor cuts the matching frames, and then the same sox-evidence-boxer + sox-evidence-context + sox-evidence-reviewer fan-out as /sox-annotate-xlsx locates, reads, and verifies each attribute, one agent per frame, in parallel, while the skill auto-judges pass/fail from their JSON.

When the evidence for a control is a meeting recording with a transcript, and you want one annotated frame per attribute the auditor verified without scrubbing the recording by hand. If the evidence is already a screenshot deck in xlsx form, use /sox-annotate-xlsx; to plan a test before reviewing evidence, use /sox-testing.

Flag Required Notes
<video.mp4> Yes The recording. Other opencv-supported formats (.mov, .mkv) also work.
<transcript> Yes .vtt / .srt (real timestamps, preferred) or .txt / .docx (timestamps estimated from word position).
--workpaper <path> No Destination. Defaults to workpapers/<control-id>/workpaper.xlsx.
--no-review No Skip the on-by-default box-position review pass on every frame.
--no-context No Skip the on-by-default value-and-context scan on every frame.
--max-review-iters <N> No Boxer revisions the review loop may request per frame (default 2, ceiling 3).

The control ID, the samples in scope, and the test attributes are usually supplied by the calling /sox-testing invocation; otherwise the skill asks once.

/sox-from-video walkthrough.mp4 walkthrough.vtt dispatches the walkthrough parser to find each (sample, test) review moment, extracts the matching frames, boxes and reads each tested field via the leaf agents, auto-judges the results, and writes per-sample-per-test detail tabs with annotated frames plus updated Summary rows.

  • Transcript quality drives frame quality. A vague transcript with no field names yields noisy timestamps: the skill asks before extracting dozens of frames.
  • Estimated timestamps drift. .txt / .docx transcripts have no real timestamps, so moments are estimated (~150 wpm) and frames can catch the auditor mid-scroll; .vtt / .srt are exact.
  • Gaps are surfaced, not swallowed. If the parser finds no frame for a (sample, test), you decide before extraction rather than silently leaving cells n/a.
  • One detail tab per (sample, test). If the auditor revisits an attribute, its frames stack in transcript order on the same tab.
  • Residual drift is flagged. Boxes the reviewer couldn’t fully resolve after the retry cap are noted in the Summary reasoning cell for a human glance.

Not audit or legal advice. Workpapers and assessments produced by these skills require review by qualified financial professionals before being relied on for SOX 404 compliance.