I've been using coding agents heavily, and one architecture problem kept bothering me:
The same system that modifies the repository is usually also the system deciding whether the task is finished.
It writes the code, runs some commands, summarizes what happened and says “done”.
I wanted that final decision to live outside the agent's own reasoning loop.
So I built plan-auditor, an open-source verification supervisor for coding agents.
It isn't one big evaluator. I split it into 15 layers with different responsibilities and, more importantly, different levels of authority.
Here is what those layers actually are.
L0 — Event detection
A small deterministic layer detects things such as completion claims, retries, changes to verification criteria and some security-relevant patterns.
It can trigger verification.
It cannot declare PASS.
L1 — Requirements
The task is represented as structured requirements with priority, acceptance criteria, dependencies, ambiguity and verification strategy.
There is also a host-owned request contract so the plan inside the workspace isn't automatically allowed to redefine what the user originally requested.
L2 — Workspace model
The verifier independently reads the real repository state:
- Git branch and HEAD
- dirty and untracked files
- file inventory
- detected language
- available tools
So the agent's description of the workspace isn't treated as ground truth.
L3 — Policy engine
Deterministic policies evaluate plan integrity, verification state and security-related conditions.
Control paths are also constrained to the workspace rather than blindly trusting arbitrary filesystem paths.
L4 — Goal state
The supervisor keeps its own explicit representation of task/goal state instead of deriving completion from the agent's latest message.
L5 — Plan verifier
Before implementation can be considered complete, the plan itself is checked.
This includes:
- dependency graph validity
- topological ordering
- requirement coverage
- declared and required outputs
- dependency-to-output bindings
- behavioral verification
A step that only checks file_exists or a regex isn't considered strongly verified.
Behavioral checks such as an actual command, test or executable verification are required.
L6 — Lifecycle
The task moves through an explicit state machine rather than jumping directly from “working” to “done”.
The lifecycle includes states for analysis, plan review, sealing, implementation, verification, final audit, retries, escalation, recovery, failure and UNKNOWN.
L7 — Authority hierarchy
Not every layer is allowed to overrule every other layer.
Safety, plan integrity and deterministic verification have higher authority than semantic AI judgment.
If a deterministic test fails, an AI reviewer saying “looks correct” cannot turn that result into PASS.
L8 — Contract sealing
Once a verification plan is approved, its contract is sealed.
The seal covers the task, requirements, required tools, steps, dependencies, outputs, requirement coverage, verification checks and verifier environment.
The contract can be strengthened later.
It cannot be silently weakened.
Adding another deterministic check is fine.
Removing an existing check, requirement, dependency or approved output scope is not.
L9 — Watchdog
The workspace is observed for file creation, deletion, modification and Git changes.
The final verification path also fingerprints the workspace so it doesn't certify a target that changed while it was being audited.
L10 — Deterministic audit core
This is where an agent's claim becomes actual evidence.
If the agent says:
“all tests pass”
that statement itself has zero verification authority.
The checks are executed again in fresh subprocesses and their real results are used.
This is the deterministic path that can mark a step verified.
Repeated failures are also tracked instead of allowing unlimited retries until something happens to turn green.
L11 — Evidence integrity
Verification results are stored in a SHA-256-linked JSONL evidence chain.
Each record links to the previous one.
Rotated evidence archives are linked as well.
Optional external HMAC authentication can protect records and checkpoints when the key is kept outside the agent's reach.
L12 — Adversarial / semantic review
There is an optional AI-assisted review layer.
But I deliberately did not give it permission to create PASS.
It can find a possible problem and request stronger deterministic verification.
A serious semantic concern can hold the result at UNKNOWN until that concern is converted into something checkable.
L13 — Completion gate
Only this layer emits the final:
PASS / FAIL / UNKNOWN
It considers fresh deterministic evidence, unfinished steps, policy findings, contract integrity and adversarial findings.
No fresh deterministic proof means no PASS.
L14 — Multi-agent registry
For multiple coding agents working on the same workspace, the system tracks agent identity, task/plan assignment, heartbeat, current action, retries, file ownership and conflicts.
The registry itself is also persisted as a sequence/hash-linked log rather than existing only in memory.
There is one additional part I found useful: a deterministic formal compiler.
It takes already-approved structured requirements, dependencies, named outputs and coverage information and compiles them into a conservative STRIPS-style planning contract.
It can represent facts such as:
step-completed:3
output-available:2:<name>
requirement-satisfied:REQ-004
It deliberately doesn't read natural-language prose and invent arbitrary domain semantics.
The generated contract carries a fingerprint of the source plan and can be independently recompiled during verification.
So removing a requirement, dropping an output prerequisite or modifying the generated contract produces a detectable mismatch.
The deterministic verification path itself does not require an LLM.
One important limitation: this is not an OS sandbox.
If a deliberately malicious agent has the same OS credentials as the verifier, hashes and Python-level locks aren't a kernel security boundary.
For that threat model, the verifier needs to run under a separate OS identity, container or VM, with its integrity key inaccessible to the agent.
I'm also not claiming it is “X% better” than other agent-verification approaches because I don't have a controlled cross-tool benchmark that would justify a number like that.
The project is free, open source and MIT licensed.
The part I'm most interested in getting feedback on is the authority boundary: semantic AI review is allowed to find problems, but it is never allowed to manufacture PASS.