Skip to content
Andy Lawsonandylawson.uk

Cloud and EUC

Enterprise AWS as code: design the state boundary first

The first useful architecture decision in an infrastructure repository is who can change which resources, through which state, and how that state will be recovered.

Andy Lawson31 July 20264 min read
Two AWS account boundaries each contain independently managed state and resources, with both consuming a shared foundation and each state having a recovery record.

A repository full of infrastructure definitions can still leave an enterprise dependent on one person's laptop. If that person holds the deployment credentials, understands the state layout and knows which resources must never be replaced, the configuration has been codified but the operating model has barely changed.

I would start an enterprise AWS design with a different question: what is the smallest useful unit of infrastructure that one accountable team can change and recover? That question connects repository structure to the business service. It also exposes decisions that a directory called production tends to conceal.

Infrastructure as Code is valuable because it makes intended configuration reviewable. The difficult work is making the relationship between that intention, the live resources and the people operating them equally explicit.

State boundaries are operational boundaries

Terraform state connects resource addresses in configuration to objects in the provider. It is therefore part of the control system, not a disposable cache. AWS explains the distinction between Terraform's configurable state storage and CloudFormation's service-managed approach in its state and backend guidance.

Putting an entire estate into one state can make initial development convenient. It can also couple unrelated changes, widen deployment permissions and make an incident in one component block work elsewhere. Splitting everything into tiny states introduces a different problem: dependencies become difficult to coordinate and outputs become an informal integration API.

Useful boundaries usually follow ownership, failure impact and change frequency. A shared network foundation may need a different release process from an application environment. A security logging account may have a different operational owner from the services sending logs to it. Separate AWS accounts help establish isolation, but an account boundary does not automatically create a sensible state boundary inside it.

For each proposed state, name the owner, authorised deployment role, consumers of its outputs and recovery procedure. If those answers are contradictory, reorganising the folders will not resolve the design.

Two AWS account boundaries each contain independently managed state and resources, with both consuming a shared foundation and each state having a recovery record.
State boundaries separate change ownership; shared dependencies remain explicit.

Protect the state as a privileged asset

State can contain sensitive material even when configuration marks an output as sensitive. Protect the backend with restricted access, encryption, auditability and an explicit recovery process. Avoid distributing copies through tickets or treating plan artefacts as safe simply because they are generated files.

For an S3 backend, HashiCorp recommends bucket versioning and documents opt-in S3 locking through use_lockfile. Its current backend documentation also marks DynamoDB locking as deprecated. That is a reason to check the installed Terraform version and plan a controlled transition, rather than copy an old backend example into a new estate.

Locking and versioning solve different problems. A lock coordinates writers. Object versions may help recover state after accidental damage. Neither proves that a recovered state matches the live environment. Restoring an old state object without checking what happened afterwards can create a misleading description of resources that still exist.

Recovery should include finding the correct version, restoring it under controlled access, refreshing the view of live resources and inspecting the proposed changes before any apply. Run that exercise in an isolated environment. A documented recovery command that nobody has tested is an assumption with good formatting.

Review the actual change, including replacements

A pull request describes code changes. A plan describes the resource actions the tooling proposes. Both deserve attention because a small source change can have a large operational consequence.

Consider a hypothetical shared network change. An engineer adjusts a module input that appears to rename a component. The resulting plan includes replacement of a resource used by several workloads. A source review concerned only with naming conventions will miss the consequence. The reviewer needs to ask what loses connectivity, which downstream states consume the identifier and whether a maintenance window is required.

The review record should connect the commit, tool and provider versions, target account, state identity and proposed actions. Deployment should execute the reviewed saved plan where supported and practical, with a fresh review if the relevant inputs or environment have changed. Protect that plan as a sensitive artefact too.

Separate the identity used to inspect an environment from the identity authorised to change it. Use short-lived credentials where the delivery platform supports them. Broad administrator access makes a pipeline easy to build, but it also makes a mistake in that pipeline much harder to contain.

Drift needs a disposition

An emergency console change is not automatically a failure of the operating model. Leaving it unexplained indefinitely is. The useful response is to establish whether the live change is legitimate, whether it should become the declared configuration and who owns that decision.

Drift reports also have coverage limits. CloudFormation's documentation explains that detection applies to supported resources and explicitly configured properties, and that nested stacks require their own checks. A reassuring status must be read alongside what the tool did not inspect.

I would attach a disposition to meaningful drift: adopt the change, restore the declaration, investigate further, or accept a time-limited exception. Automatically reverting every difference can undo an incident mitigation. Automatically adopting every difference turns the repository into a transcript of whatever happened in the console.

The distinction is the same one that matters in evidence-led control changes: observation supports a decision; it does not make the decision for you.

Introduce code without recreating the estate

Existing AWS resources should enter management through an inventory and a controlled adoption process. Resource import associates existing objects with declared addresses; it does not prove that the declaration fully represents their configuration or operational purpose.

Start with a bounded component. Compare the imported resource against live configuration, inspect the first plan and resolve unexpected changes before expanding scope. Keep a record of unmanaged dependencies. A clean plan for one module says nothing about a certificate, route or scheduled process maintained elsewhere.

The completion criterion is operational: another authorised engineer can explain the state boundary, review a change, run the deployment and recover its control data. At that point the repository has become a dependable part of the platform, rather than another place where specialist knowledge is hidden.