How I Learned the Hard Way That Terraform Needs Policy as Code

At a glance
- Context: a cloud consultancy engagement, bootstrapping 15 new AWS accounts for a client with Terraform 0.11/0.12.
- Incident: I accidentally deleted the S3 bucket that held the Terraform state for all 15 accounts.
- Cost: all 15 accounts were recreated and provisioned again from scratch. Because everything was code, that took less than two days. The real cost was our reputation with the client.
- Fix: policy as code. OPA policies evaluated against every
terraform planin CI, failing the pipeline before a dangerous change reachesapply. - Outcome: 15 workload pipelines gated, and every pipeline created after them inherited the same checks.
What went wrong?
A few years ago, while working for a cloud consultancy company, I made a mistake that I will never forget. I accidentally deleted a Terraform state bucket: the S3 bucket that stored the state for bootstrapping 15 AWS accounts, each packed with resources. To make things worse, it belonged to a very tough client.
If you’ve worked with Terraform, you know how critical the state file is. Without it, Terraform has no record of what it has already deployed. My mind raced through options. I considered importing every resource back into state by hand. We even contacted AWS Support to ask whether the bucket could be restored. In the end, the only real option was to tell the client.
We were lucky in one way: the accounts were brand new and had no live workloads yet. After some tense discussions with a very pushy client, we agreed to start over, creating 15 new accounts and provisioning everything again from scratch. Because everything was infrastructure as code, that took less than two days. The expensive part was the damage to our reputation with the client.
The deletion was the trigger, not the root cause. The root cause was that nothing stood between a person and a destructive change. Our only safety net was an engineer reading the variables files before running apply.
Why wasn’t infrastructure as code enough?
Infrastructure as code made the recovery fast. It did nothing to prevent the mistake. Terraform makes changes repeatable, including the bad ones, and it has no opinion on whether a change is safe.
Manual review worked while our code was simple: a series of module invocations, with any risky setting in plain sight in the tfvars files. As the deployments grew, with nested modules and autogenerated tfvars files, a dangerous value could be set several layers away from anything a reviewer would open. Catching it by eye stopped being realistic.
What I needed was a rule that runs automatically, on every change, against what Terraform is actually about to do. That is the plan. terraform show -json turns a saved plan into a JSON document with every value that is known at plan time resolved through all the modules and variable files, which is exactly what a reviewer couldn’t see.
At the time, around Terraform 0.11 and 0.12, there weren’t many tools for validating plan output. The one I found was Open Policy Agent (OPA): a general-purpose policy engine that evaluates rules, written in a language called Rego, against any JSON document. A Terraform plan is just another JSON document.
How does the guardrail work?
The pipeline gets one extra step between plan and apply:
terraform plan -out=tfplan.binary
terraform show -json tfplan.binary > tfplan.json
conftest test tfplan.json -p policies/terraform --all-namespaces
Conftest evaluates the OPA policies in policies/terraform against the plan and exits non-zero on any violation, so the pipeline stops before apply runs.
--all-namespacesmatters. By default Conftest only evaluates themainpackage. A policy in any other package, liketerraform.planbelow, is silently skipped: the step prints0 testsand passes. A guardrail that doesn’t run is worse than none, because everyone believes it does.
The first policy targets the setting that turns deleting a bucket into a single, irreversible step. force_destroy = true lets Terraform delete an S3 bucket together with everything in it:
package terraform.plan
deny contains msg if {
some resource in input.resource_changes
resource.type == "aws_s3_bucket"
resource.change.after.force_destroy == true
msg := sprintf("S3 bucket '%s' has force_destroy = true: one apply can delete it with everything in it", [resource.address])
}
Against a plan that sets it, the pipeline fails and names the resource:
FAIL - tfplan.json - terraform.plan - S3 bucket 'aws_s3_bucket.state' has force_destroy = true: one apply can delete it with everything in it
1 test, 0 passed, 0 warnings, 1 failure, 0 exceptions
The policy uses Rego v1 syntax (deny contains msg if), which OPA 1.0 and later require. Tested with Conftest (OPA 1.20.2) and Terraform 1.16.3.
The rest of the series builds this out step by step:
- First policy as code using OPA: writing and testing this policy with the OPA CLI and Conftest.
- Enforcing AWS resource tags with OPA: required tags and allowed values, with the data kept outside the policies.
- Enforcing AWS AMI images with OPA: only approved AMIs make it to
apply. - Using Shift Left paradigm to manage cloud cost: cost guardrails with OPA and Terracost.
What changed?
The class of mistake that cost us 15 accounts is now caught at plan time, with the offending resource named in the pipeline log, before anyone can run apply. Safety no longer depends on a reviewer spotting one value in a generated file.
Once the pipeline step existed, each new rule became cheap to add. Tagging standards, approved AMIs and cost limits followed as more policies evaluated in the same step.
The policies gated 15 workload pipelines, and every new pipeline inherited the check without extra work. It stopped dangerous plans many times after that, including in pipelines nobody had to remember to protect.
Where should a team start?
- Start with your most expensive mistake. One policy that prevents it is worth more than fifty you don’t need yet.
- Warn before you deny. Write a new rule as
warnfirst: Conftest reports the violation but the pipeline still passes. Clean up what it finds, then switch the rule todeny. - Check the check. A healthy run reports at least one test (
1 test, 1 passed).0 testsmeans your policies aren’t being evaluated at all.
If your Terraform pipeline goes straight from plan to apply with nothing in between, that’s the kind of gap I help teams close as a fractional DevOps engineer. Get in touch.


