Skip to content
Andy Lawsonandylawson.uk

AI-assisted engineering

Agentic development needs an independent verification boundary

An agent's implementation, tests and completion report can all share the same mistaken assumption. Verification needs a reference outside that loop.

Andy Lawson4 September 20264 min read
Independently established expected results and observed results from a bounded generated implementation meet at a comparison gate before acceptance.

A coding agent can produce an implementation, write tests for it, run those tests and explain why the work is complete. That is a useful delivery loop. It is also a loop in which one mistaken assumption can survive every stage.

Suppose a hypothetical administrative action should affect only records belonging to the current organisation. The agent implements a filter using the wrong organisation identifier, then constructs test fixtures using that same identifier. The suite passes. The explanation is coherent. The isolation requirement has still been misunderstood.

The answer is an independent verification boundary: expected behaviour, evidence or authority that is not derived solely from the generated implementation. Independence can come from a specification, a known dataset, an existing contract or a separately designed test. It does not require every check to be manual.

Establish the expected result before the code

For consequential logic, choose examples with externally established outcomes. If the feature calculates a result, use inputs whose correct answer has been checked independently. If it enforces access, define the allowed and denied cases from the authorisation model rather than the interface design.

Include boundary conditions. A date rule should cover the exact transition and the relevant time zone. An import should cover duplicate identifiers and malformed records. A tenant boundary should cover attempts to fetch, update and delete another tenant's objects.

Keep the expected result stable while the implementation changes. Otherwise the agent can accidentally resolve a failing test by changing the test's expectation to match the defect. When a requirement genuinely changes, record that decision explicitly.

NIST's SSDF provides an established framework for integrating security practices into software development. The agent-specific implication is practical: verification obligations belong to the delivery process, regardless of who or what generates the code.

Independently established expected results and observed results from a bounded generated implementation meet at a comparison gate before acceptance.
Agreement inside the generation loop is insufficient; compare against an independent reference.

Test at the boundary where enforcement happens

Interface checks matter, but a hidden control is not an authorisation mechanism. An unauthorised user should also be denied when making the equivalent request directly to the server.

Test the relevant API using authenticated identities with different rights. Verify both the response and the persisted state. A denied response that arrives after a mutation has already occurred does not satisfy the requirement.

For a data workflow, trace representative input through parsing, transformation, persistence and presentation. A unit test may establish that a parser handles a sample correctly while the real upload route passes it the wrong encoding or strips required context. The complete journey needs its own evidence.

Do not make every end-to-end test enormous. Use a small number of deliberate journeys that cross important boundaries, supported by focused tests for individual rules. The aim is useful coverage of failure modes, not a large suite that mostly repeats the implementation.

Limit the agent's authority

Verification becomes harder when the tool producing the change can also modify the environment that defines correctness. An agent with unrestricted access to source, tests, deployment credentials and production data has a much wider failure surface than one working in a bounded development environment.

Give it the capabilities required for the assignment. Keep production credentials and sensitive datasets outside that boundary unless there is a clear, authorised need. A development token should have development permissions, not merely a reassuring filename.

Claude Code's security documentation distinguishes permission controls from operating-system sandboxing and describes protections against prompt injection. Those mechanisms are useful, but their presence does not establish that a particular session has safe permissions. Check the configured mode, accessible paths, network destinations and available credentials.

An isolated checkout protects concurrent source changes. It does not, by itself, isolate network access or secrets inherited from the process environment. Treat those as separate controls.

Treat retrieved material as untrusted input

Agents read repositories, issue descriptions, documentation and tool responses. Any of those can contain text that looks like an instruction. The material may be useful evidence without being authorised to redefine the task.

For example, an issue body could tell an agent to disable a security check before running a test. A dependency README could recommend a command that sends environment data to an external service. The engineer should evaluate the action and its source rather than accept the imperative because it appeared in context.

Reduce the opportunity for that confusion. Keep project instructions explicit. Use narrow tool permissions, inspect unfamiliar commands and avoid loading unnecessary sensitive content. Logging helps reconstruct actions, but prevention depends on the capability boundary being real.

Security review should include generated dependency changes and install scripts. A feature that works can still introduce an unnecessary package, broader access or a new external data path. Those are design changes and deserve the same scrutiny as the feature itself.

Require a reproducible completion account

A useful handover names the requirement, the evidence that verifies it and the exact revision tested. It distinguishes checks that passed from checks that were not run. It explains relevant environmental limitations without turning them into a claim of success.

Ask for failures discovered and corrected where they affect confidence. A short account of a boundary defect and its verification can be more valuable than a long list of green commands. The reviewer needs to know why the final state meets the contract.

The existing article on control-change evidence asks whether an on-call engineer could reconstruct a consequential change. Apply the same standard here: another engineer should be able to reproduce the verification without trusting the agent's narrative.

Keep acceptance outside the generation loop

An agent may propose that the work is ready. The accountable engineer decides whether the evidence is sufficient for the intended deployment. That decision can be quick for a bounded, low-impact change and more demanding for identity, financial logic or destructive operations.

The boundary should scale with consequence. It should not become a ceremony applied identically to every edit. Its purpose is to preserve an independent account of correctness while making good use of automation. Agentic development becomes dependable when generated work can be accepted on evidence that remains meaningful outside the agent's own loop.