Skip to content
Andy Lawsonandylawson.uk

AI-assisted engineering

AI-assisted engineering starts with an acceptance contract

A coding agent needs a precise definition of the behaviour it must deliver. Without that, faster implementation can simply create a larger review queue.

Andy Lawson14 August 20265 min read
An acceptance contract defines implementation; evidence returns through review to be compared with the independently stated expected result.

“Add a better upload workflow” sounds like a reasonable request until somebody has to decide whether the result is finished. Does better mean faster, more informative, less likely to lose a selection, or able to recover after an interrupted connection? A coding agent can produce a convincing interpretation of any of those possibilities.

The expensive part comes afterwards, when an engineer reviews a polished implementation and discovers that the intended behaviour was never agreed. The conversation becomes a sequence of corrections, each applied to code that already contains assumptions from the previous version.

My starting point is an acceptance contract: a short, explicit description of observable behaviour, boundaries and evidence. It is not a legal document or a large requirements exercise. It is a way to make the assignment testable before generation begins.

Describe the journey in concrete terms

For a hypothetical upload feature, the contract might say that selecting a second group of files adds them to the current selection; it does not replace it. Removing one file affects only that file. Refreshing during an upload must have a documented outcome. A failed file must remain identifiable, with a retry path that does not duplicate successful work.

That description is already more useful than a list of components. It gives the implementer freedom over the interface while fixing the behaviour that matters to the user.

A compact contract makes the checks explicit:

  • A second selection adds to the first.
  • Removing one file leaves every other selected file in place.
  • Retrying a failed file does not duplicate completed work.
  • A refresh has a defined outcome that the interface explains.

Add the relevant constraints: supported file types, size handling, authorisation, accessibility and where data is stored. Distinguish decisions already made from decisions the agent may propose. Otherwise a small interaction change can quietly expand into a new storage design or an unnecessary dependency.

I would also state what existing behaviour must survive. An enhancement is incomplete if it works in isolation but breaks the previous workflow. That requirement is particularly important when an agent sees only part of a mature repository.

Ask for evidence before asking for polish

An acceptance contract should define how success will be demonstrated. For the upload example, an automated interaction check can verify additive selection. An integration test can verify that a retry does not create another stored object. A manual check can assess whether the error state makes sense to a user.

Those checks answer different questions. A screenshot proves appearance at one moment. A test that mocks every external boundary proves an internal interpretation. Neither demonstrates that the complete workflow behaves correctly against the real application interface.

The contract should therefore identify the smallest complete journey that must be exercised. That may be browser, API, storage and response, or a command-line operation followed by inspection of its persisted result. The verification must cross the boundary where the requirement can actually fail.

This is consistent with the broader discipline in NIST's Secure Software Development Framework: security practices belong throughout the development lifecycle. Adding an agent does not remove the need to define and verify the behaviour being shipped.

An acceptance contract defines implementation; evidence returns through review to be compared with the independently stated expected result.
The requirement supplies the expected result; implementation supplies the evidence.

Keep the work unit reviewable

An agent can edit more files in a sitting than a human can comfortably review. That capability makes scope control more valuable. A request that combines a new workflow, architectural refactoring and visual redesign can leave the reviewer unable to isolate which change caused a regression.

Choose a work unit that delivers useful behaviour and has a clear verification boundary. It need not be tiny. A complete feature spanning several layers can be easier to assess than a series of disconnected fragments, provided its purpose and acceptance evidence stay coherent.

The agent's handover should explain the resulting behaviour, material design decisions and tests actually run. It should identify incomplete checks and assumptions that still matter. A list of modified files is helpful navigation, but it is not an engineering account of the change.

When a contract reveals genuine ambiguity, resolve it before implementing the dependent part. Meanwhile, the agent can inspect the existing architecture or prepare checks that do not depend on that decision. Autonomy is useful when it advances understood work; it is costly when it multiplies assumptions.

Separate implementation from judgement

The acceptance contract should come from the requirement and the system's obligations. It should not be reverse-engineered from whatever the agent happened to produce.

That matters for tests. If the same generation process creates both an implementation and a test that copies its logic, they can agree perfectly while misunderstanding the requirement. Independent examples help: a known valid input, an invalid boundary case and an expected outcome established before the code exists.

For access control, test what an unauthorised user can attempt through the API, not just whether the interface hides the button. For data transformations, include known awkward inputs with expected outputs checked separately. For destructive operations, verify the affected set and recovery behaviour.

The engineer remains responsible for deciding whether the evidence is enough. Delegating code generation does not transfer the authority to accept risk or to redefine the business rule when implementation is inconvenient.

Measure the whole delivery loop

The useful productivity measure is the effort needed to reach accepted, supportable behaviour. Include requirement clarification, generation, review, correction, integration and verification. Generated lines, completed prompts and first-pass speed can all improve while that total becomes worse.

METR's study of experienced open-source developers using early-2025 tools found that perceived benefit and measured completion time could diverge. Its particular setting should not be treated as a forecast for every tool or team in 2026. It does justify measuring the actual workflow instead of relying on how productive a session feels.

Record where time goes. A team may discover that the agent helps with implementation but requirements remain the bottleneck. Another may find that a reviewer spends too long explaining local conventions. Those are actionable findings: improve the contract, repository guidance or validation harness, then compare again.

The existing article on where AI helps in platform engineering deals with coverage and human accountability. An acceptance contract applies that principle to software delivery: define the outcome, bound the change and require evidence that survives independent inspection. That is how a fast coding session becomes completed engineering work.