Software Engineering WikiSE Wiki

Terraform

Plan and apply safely, structure modules and variables, refactor and repair state, and recover from drift, lock and provider errors in Terraform.

Reviewed MarkdownEdit

On this page

Cheatsheet#

TaskCommand
Initialise backend and download providersterraform init
Move state to a changed backendterraform init -migrate-state
Upgrade providers within constraintsterraform init -upgrade
Format and validateterraform fmt -recursive && terraform validate
Plan to a fileterraform plan -out=tfplan
Apply exactly that planterraform apply tfplan
Detect drift without proposing changesterraform plan -refresh-only
Recreate one resourceterraform apply -replace='aws_instance.web[0]'
List resources in stateterraform state list
Inspect one resource in stateterraform state show 'aws_instance.web[0]'
Back up remote stateterraform state pull > backup.tfstate
Release a stale lockterraform force-unlock <lock-id>
Output for scriptsterraform output -raw url
Run module teststerraform test

Current release at review time: Terraform 1.16. Version-dependent features below state the version that introduced them. See the Terraform documentation.

How plan and apply work#

Terraform reads the configuration, builds a dependency graph from references between blocks, refreshes state by asking providers for the current attributes of every tracked resource, then diffs desired against current and proposes create, update, replace or destroy actions. Independent resources are applied in parallel (10 at a time by default, -parallelism=n).

State maps each resource address (aws_instance.web[0]) to a real object ID. It is the only link between configuration and the real world: a resource missing from state is invisible to Terraform, and a resource removed from configuration but still in state is scheduled for destruction.

Three things can disagree: configuration, state and reality.

CommandComparesChanges infrastructure
terraform planConfiguration against refreshed stateNo
terraform plan -refresh-onlyState against reality (drift)No
terraform apply -refresh-onlyWrites reality into stateNo, state only
terraform applyConfiguration against refreshed stateYes
terraform plan -out=tfplan
terraform show tfplan                   # human-readable review of the saved plan
terraform apply tfplan                  # applies exactly what was reviewed, or fails if state moved on

Always use plan -out then apply <file> in automation. apply without a plan file re-plans, so what runs can differ from what was reviewed. A saved plan fails to apply if the state changed after it was created.

Plan symbols: + create, - destroy, ~ update in place, -/+ destroy then create, +/- create then destroy (create_before_destroy), <= read a data source. Look for # forces replacement next to an attribute to see why a resource is being replaced.

Configuration and providers#

terraform {
  required_version = ">= 1.11"
  required_providers {
    aws = { source = "hashicorp/aws", version = "~> 6.0" }
  }
  backend "s3" {
    bucket       = "example-tfstate"
    key          = "prod/network/terraform.tfstate"
    region       = "ap-southeast-2"
    use_lockfile = true      # S3-native locking; DynamoDB locking is deprecated
    encrypt      = true
  }
}

provider "aws" {
  region = "ap-southeast-2"
}

provider "aws" {
  alias  = "us_east_1"        # for resources that must live in us-east-1, such as CloudFront certificates
  region = "us-east-1"
}

terraform init writes .terraform.lock.hcl with the exact provider versions and checksums selected. Commit it: it is what makes two machines use the same provider build. init -upgrade moves to the newest version the constraints allow and rewrites the lock file.

The dynamodb_table backend argument still works but is deprecated in favour of use_lockfile; both can be set during migration. See the S3 backend.

Resources and meta-arguments#

resource "aws_instance" "web" {
  for_each               = var.web_nodes          # map of name => { subnet_id = ... }
  ami                    = data.aws_ami.al2023.id
  instance_type          = var.instance_type
  subnet_id              = each.value.subnet_id
  vpc_security_group_ids = [aws_security_group.web.id]

  tags = merge(var.tags, { Name = "${var.name}-${each.key}" })

  lifecycle {
    create_before_destroy = true
    ignore_changes        = [ami]          # an image pipeline owns this attribute
    precondition {
      condition     = var.instance_type != "t2.micro"
      error_message = "t2.micro is not permitted in production."
    }
  }
}
Meta-argumentEffect
countInstances indexed by position (web[0]); removing a middle element shifts every later index
for_eachInstances keyed by map key or set value (web["a"]); adding or removing a key touches only that key
depends_onExplicit ordering when no expression reference exists
providerSelects an aliased provider configuration (another region or account)
lifecycle.create_before_destroyCreates the replacement before destroying the old object
lifecycle.prevent_destroyFails any plan that would destroy the resource; use on databases and state buckets
lifecycle.ignore_changesStops Terraform reverting changes to listed attributes
lifecycle.replace_triggered_byReplaces this resource when a referenced resource or attribute changes

Prefer for_each for anything with a natural key. With count, deleting the first list element destroys and recreates every resource after it. for_each keys must be known at plan time, so they cannot come from attributes of resources that do not exist yet.

Variables, outputs and precedence#

variable "instance_type" {
  type        = string
  default     = "t3.small"
  description = "EC2 instance size for the web tier"
  validation {
    condition     = can(regex("^t3\\.", var.instance_type))
    error_message = "Only t3 sizes are approved here."
  }
}

output "url" {
  value       = "https://${aws_lb.this.dns_name}"
  description = "Public endpoint"
}

Precedence, lowest to highest: the variable’s default, TF_VAR_<name> environment variables, terraform.tfvars, terraform.tfvars.json, *.auto.tfvars and *.auto.tfvars.json in lexical order, then -var and -var-file in command-line order. A later source replaces an earlier value entirely; maps are not merged.

Keeping secrets out of state#

sensitive = true hides a value from CLI output only. It is still written in plain text to state and saved plan files.

FeatureVersionStored in state or plan
sensitive = true on variables and outputs0.15Yes, redacted in output only
ephemeral = true on variables and outputs, ephemeral resource blocks1.10No
Write-only resource arguments (usually named *_wo, paired with a *_wo_version)1.11No
ephemeral "random_password" "db" {
  length = 32
}

resource "aws_db_instance" "main" {
  # ...
  password_wo         = ephemeral.random_password.db.result
  password_wo_version = 1          # bump to push a new password
}

Write-only arguments exist only where the provider implements them; check the resource documentation. See managing sensitive data and, for issuing secrets at apply time, Vault.

State is a secret

Encrypt the backend, restrict read access to the people and pipelines that run Terraform, turn on bucket versioning so a bad write can be rolled back, and never commit terraform.tfstate or *.tfvars files holding credentials.

Modules#

module "network" {
  source = "git::https://github.com/example/tf-modules.git//network?ref=v1.4.0"

  name     = "prod"
  cidr     = "10.20.0.0/16"
  az_count = 3
}

module "vpc" {
  source  = "terraform-aws-modules/vpc/aws"     # registry module
  version = "~> 6.0"
  # ...
}

Pin module sources to a tag or commit (?ref=) or a registry version. A floating branch lets a plan change because someone else merged. Run terraform init (or init -upgrade) after changing a source or version.

Keep modules to inputs, resources and outputs. A module that hard-codes naming, tagging or environment decisions is harder to reuse than one that accepts them as variables. terraform test (1.6+) runs *.tftest.hcl files that plan or apply a module and assert on the result.

Refactoring and repairing state#

Prefer configuration blocks to state commands: they appear in the plan, are reviewed in a pull request and apply the same way in every workspace.

GoalDeclarative (reviewed in plan)Imperative (immediate, unreviewed)
Rename or move into a modulemoved block (1.1+)terraform state mv
Stop managing, keep the objectremoved block with lifecycle { destroy = false } (1.7+)terraform state rm
Adopt an existing objectimport block (1.5+; for_each 1.7+)terraform import
moved {
  from = aws_instance.web
  to   = module.compute.aws_instance.web
}

removed {
  from = aws_instance.legacy
  lifecycle {
    destroy = false
  }
}

import {
  to = aws_s3_bucket.logs
  id = "example-logs-bucket"
}
terraform plan -generate-config-out=generated.tf   # writes resource blocks for import targets that lack one

Review generated configuration before applying; it contains every attribute, including defaults you should delete.

Back up before state surgery

Run terraform state pull > "state-$(date +%Y%m%dT%H%M%S).json" before any state mv, state rm, state push or force-unlock. These commands change remote state immediately with no plan.

terraform state list
terraform state show 'aws_instance.web["a"]'
terraform state mv 'aws_instance.web' 'aws_instance.api'
terraform state rm 'aws_instance.legacy'            # forget it; the object keeps running
terraform import 'aws_instance.web["a"]' i-0abc123

Quote addresses that contain [ or " so the shell passes them unchanged.

moved, import and removed in depth#

moved blocks chain: when a resource is renamed twice across releases, keep both blocks so any state at either old address migrates. They work across module boundaries and for whole modules, and for count to for_each conversions where each index needs its own block.

moved {                                   # whole module rename; every resource inside moves with it
  from = module.net
  to   = module.network
}

moved {                                   # count to for_each: one block per index
  from = aws_subnet.private[0]
  to   = aws_subnet.private["a"]
}
moved {
  from = aws_subnet.private[1]
  to   = aws_subnet.private["b"]
}

import {                                  # bulk import driven by a map (1.7+)
  for_each = var.existing_buckets          # { logs = "example-logs", assets = "example-assets" }
  to       = aws_s3_bucket.managed[each.key]
  id       = each.value
}

import {
  to       = aws_route53_record.www
  id       = "Z0123456789ABC_www.example.com_A"     # ID formats are provider-specific; check the resource docs "Import" section
}                                                   # 1.12+: providers may accept an identity {...} object instead of an id string

removed {
  from = module.legacy                    # also works for whole modules
  lifecycle { destroy = false }
}

A plan shows # (moved from ...) and # (imported from ...) lines, and an import for an object whose attributes differ from the configuration proposes an update in the same plan; read that update before applying. Once applied, delete the import and moved blocks in a later change (they are harmless while present, but moved blocks whose from no longer exists in anyone’s state are noise). terraform plan -generate-config-out refuses to overwrite an existing file and only generates for import blocks whose target has no configuration.

State surgery#

State is JSON with a serial and lineage. The backend rejects a push whose serial is not newer or whose lineage differs, which is the safety net when hand-editing. Every operation below rewrites remote state immediately.

terraform state pull > state.json                          # always first; keep this copy until the next successful plan
terraform state list -id i-0abc123                         # which address holds an object with that ID
terraform state show -no-color 'module.db.aws_db_instance.main' | grep -E '^\s+(arn|id) '
terraform state mv 'module.old.aws_instance.web' 'module.new.aws_instance.web'
terraform state rm 'module.legacy'                         # forgets every resource under the module
terraform state replace-provider registry.terraform.io/-/aws registry.terraform.io/hashicorp/aws   # after the 0.13 provider namespace change or a fork
terraform apply -replace='aws_instance.web["a"]'           # supersedes terraform taint
terraform untaint 'aws_instance.web["a"]'                  # clear a taint left by a failed create
terraform apply -refresh-only                              # accept reality into state without touching configuration
terraform state push state.json                            # upload an edited copy; refuses older serial or different lineage
terraform state push -force state.json                     # overrides both checks; last resort

Moving a resource between two root modules (two backends) has no single command. Pull both states, move with the local-file form, then push both:

cd ../network && terraform state pull > /tmp/net.json && cd ../compute && terraform state pull > /tmp/comp.json
terraform state mv -state=/tmp/net.json -state-out=/tmp/comp.json 'aws_security_group.web' 'aws_security_group.web'
cd ../network && terraform state push /tmp/net.json && cd ../compute && terraform state push /tmp/comp.json
# then add the resource block to compute, delete it from network, and plan both: each should show no changes

Editing the JSON by hand is occasionally necessary (a provider bug wrote an impossible attribute, or a resource must be nudged to a new schema version). Bump serial, keep lineage, and validate with terraform show state.json before pushing. A lost state file is recovered by importing every object, which is why terraform state list output belongs in the run logs of any pipeline.

Workspaces vs directories#

Two ways to run one configuration against several environments:

CLI workspacesOne directory per environment
StateSame backend, separate state per workspace (env:/<name>/ prefix on S3)Separate backend configuration and state per directory
CredentialsShared; the same principal can touch every environmentCan differ per directory (separate roles, accounts, subscriptions)
Configuration drift between environmentsImpossible: one set of filesPossible; controlled by sharing modules and diffing tfvars
Blast radius of a wrong applyHigh: terraform workspace select is easy to forgetLow: -chdir=envs/prod is explicit
SuitsEphemeral copies (review environments, per-branch stacks)Long-lived environments with different sizes, providers and approvers
envs/
  prod/    main.tf  backend.tf  prod.tfvars      # root module: a handful of module calls and provider config
  staging/ main.tf  backend.tf  staging.tfvars
modules/
  network/ compute/ database/                     # all environment differences arrive through variables
terraform -chdir=envs/prod init -backend-config=prod.s3.tfbackend   # partial backend config kept out of source when it holds account-specific values
terraform -chdir=envs/prod plan -var-file=prod.tfvars -out=tfplan
terraform workspace new pr-1234 && terraform apply -var-file=review.tfvars -auto-approve   # ephemeral copy
terraform workspace select default && terraform workspace delete pr-1234                  # refuses while its state is non-empty; destroy first

terraform.workspace in expressions (name = "app-${terraform.workspace}") makes workspaces usable but also hides the environment from the reader of the configuration. If a configuration needs count = terraform.workspace == "prod" ? 3 : 1, it wants directories.

Tests#

terraform test runs *.tftest.hcl files in the root or tests/ directory. Each run block plans (command = plan) or applies (command = apply, the default) the module with the given variables and checks assert conditions; applied resources are destroyed when the file finishes. Mock providers (1.7+) return fabricated values so unit tests need no credentials.

# tests/network.tftest.hcl
variables {
  name = "test"
  cidr = "10.99.0.0/16"
}

mock_provider "aws" {}                     # every resource and data source returns generated values

run "creates_one_subnet_per_az" {
  command = plan
  variables { az_count = 3 }
  assert {
    condition     = length(aws_subnet.private) == 3
    error_message = "expected 3 private subnets, got ${length(aws_subnet.private)}"
  }
}

run "rejects_small_cidr" {
  command = plan
  variables { cidr = "10.99.0.0/28" }
  expect_failures = [var.cidr]             # the variable's validation block must reject it
}

run "outputs_match" {
  command = plan
  assert {
    condition     = output.vpc_cidr == var.cidr
    error_message = "output must echo the input CIDR"
  }
}
terraform test                                   # every test file
terraform test -filter=tests/network.tftest.hcl  # one file
terraform test -verbose                          # print the plan or state for each run
terraform test -junit-xml=report.xml             # JUnit report for CI

Use command = plan with mocks for fast unit checks of logic (counts, names, validation), and a small number of real apply runs against a sandbox account for integration. A run block can reference outputs of an earlier run (run.setup.some_output) and can load a helper module with module { source = "./tests/setup" } to create prerequisites. override_resource and override_data blocks replace one object’s attributes without mocking the whole provider.

Provider lock and installation#

.terraform.lock.hcl records, per provider, the selected version, the constraints in effect and checksums (h1: for the zip on this platform, zh: for every platform’s zip from the registry). terraform init on a platform whose h1: hash is missing fails with a checksum error unless the lock file was generated for that platform too.

terraform providers                                                # required providers and which module requires them
terraform providers lock -platform=linux_amd64 -platform=darwin_arm64 -platform=linux_arm64   # record hashes for every platform CI and laptops use
terraform init -upgrade                                            # newest versions within constraints; rewrites the lock file
terraform providers mirror ./mirror                                # download providers for an air-gapped or rate-limited environment
terraform providers schema -json | jq '.provider_schemas | keys'   # inspect installed provider schemas
# ~/.terraformrc: reuse downloaded providers across working directories and prefer a local mirror
plugin_cache_dir = "$HOME/.terraform.d/plugin-cache"
provider_installation {
  filesystem_mirror { path = "/opt/terraform/mirror", include = ["registry.terraform.io/hashicorp/*"] }
  direct { exclude = ["registry.terraform.io/hashicorp/*"] }
}

TF_PLUGIN_CACHE_DIR does the same as plugin_cache_dir from the environment. The cache is not safe for concurrent init runs of different working directories on the same host (a known limitation); serialise them in CI or use a mirror. Version constraints belong in required_providers of the root module; child modules should state only the minimum they need (>= 5.0) so the root decides.

Expressions worth knowing#

locals {
  env_tags = merge(var.tags, { Environment = var.env, ManagedBy = "terraform" })
  subnets  = { for az in var.azs : az => cidrsubnet(var.cidr, 8, index(var.azs, az)) }   # for expression over a list into a map
  public   = [for s in aws_subnet.all : s.id if s.tags["tier"] == "public"]             # filter
  cfg      = yamldecode(file("${path.module}/config.yaml"))
  name     = coalesce(var.name_override, "${var.project}-${var.env}")
}

dynamic "ingress" {                       # repeat a nested block per element
  for_each = var.ingress_rules
  content {
    from_port   = ingress.value.port
    to_port     = ingress.value.port
    protocol    = "tcp"
    cidr_blocks = ingress.value.cidrs
  }
}

check "endpoint_answers" {                # post-apply assertion that does not block the apply (1.5+)
  data "http" "health" { url = "https://${aws_lb.this.dns_name}/health" }
  assert {
    condition     = data.http.health.status_code == 200
    error_message = "health endpoint returned ${data.http.health.status_code}"
  }
}
terraform console <<< 'cidrsubnet("10.0.0.0/16", 8, 3)'   # evaluate expressions against the current state
terraform console <<< 'keys(module.network.subnet_ids)'

try(expr, fallback) and can(expr) handle optional attributes; one() turns a zero-or-one element list into a value or null; sensitive() and nonsensitive() adjust redaction. templatefile() renders a file with variables and is preferable to long heredocs for user data and policies.

Workspaces and environment layout#

CLI workspaces keep separate state files under one backend configuration (for S3, under the env:/ key prefix). They suit short-lived copies of identical infrastructure, such as a feature branch stack. Production and development usually differ in size, providers and credentials, so give them separate root directories with separate backends and credentials instead.

terraform workspace list
terraform workspace new feature-x
terraform workspace select default

terraform.workspace holds the current name for use in expressions.

Troubleshooting#

SymptomCauseFix
Error acquiring the state lockAnother run is active, or a crashed run left the lockConfirm no run is active (CI, colleagues), then terraform force-unlock <id>
Saved plan is staleState changed after the plan was writtenPlan again
Plan replaces many resources after a list editcount index shiftSwitch to for_each and add moved blocks
Plan replaces resources after a provider upgradeChanged defaults or schema in the new major versionRead the provider upgrade guide; pin the old version until handled
Resource exists but plan wants to create itNot in state (created by hand, or state lost)Import it
Resource gone but plan wants to update itDeleted outside Terraformterraform apply -refresh-only, then plan
Perpetual diff on every planAPI normalises the value (case, JSON ordering, defaults)Match the normalised form, or ignore_changes
Provider produced inconsistent result after applyProvider bug or eventually consistent APIRe-run; report upstream if it persists
Invalid for_each argument ... will be known only after applyKeys depend on unknown valuesBuild keys from variables or static values
Cycle: errorTwo resources reference each other, often security groupsSplit rules into separate rule resources
Inconsistent dependency lock fileProvider added without re-running initterraform init (or init -upgrade) and commit the lock file
Slow plansThousands of resources in one state; refresh calls every APISplit state by lifecycle and blast radius; -refresh=false for a quick look only
Failed to install provider ... checksum mismatch or doesn't match any of the checksumsLock file has hashes for another platform onlyterraform providers lock -platform=... for every platform, commit the lock file
Backend configuration changedBackend block differs from .terraform/terraform.tfstateterraform init -migrate-state to move, or -reconfigure to point elsewhere without copying
import block plans an update as well as the importConfiguration differs from the real object’s attributesAlign the configuration with terraform state show output after import, or accept the change knowingly
Moved object still exists at ... from addressBoth from and to addresses exist in stateRemove one with state rm after confirming which is real
terraform test fails with credentials errorsrun blocks default to command = apply against real providerscommand = plan plus mock_provider, or supply sandbox credentials
workspace delete refusesWorkspace state is not emptyterraform destroy in that workspace first, or -force to abandon the objects (they keep running)
Error: Unsupported attribute after a provider upgradeAttribute renamed or removed in the new major versionProvider changelog and upgrade guide; terraform providers schema -json to see the current schema
Plan wants to destroy everythingWrong workspace, wrong backend key, or empty state after a failed migrationterraform workspace show, terraform state list, stop and compare with the backup
Error: Invalid function argument on file()Path relative to the working directory, not the moduleUse ${path.module}/file
ephemeral value used in a non-ephemeral contextEphemeral values may only feed write-only arguments, provider config or other ephemeralsRestructure; do not copy the value into a normal attribute
TF_LOG=DEBUG terraform plan 2>debug.log      # core and provider logs; may contain secrets
TF_LOG_PROVIDER=TRACE terraform apply        # provider API calls only
terraform providers                          # provider requirements per module
terraform graph | dot -Tsvg > graph.svg      # needs graphviz

-target is for recovery

-target plans a subset of the graph and can leave state inconsistent with configuration. Use it to get out of a broken state, then run a full plan.

Oneliners#

# Every change in a saved plan, one line each
terraform show -json tfplan | jq -r '.resource_changes[] | select(.change.actions != ["no-op"]) | "\(.change.actions|join(","))\t\(.address)"'

# Only the deletes and replacements: the review that matters
terraform show -json tfplan | jq -r '.resource_changes[] | select(.change.actions | index("delete")) | .address'

# Resource counts by type
terraform state list | sed 's/\[.*//' | awk -F. '{print $(NF-1)}' | sort | uniq -c | sort -rn

# Outputs into environment variables
eval "$(terraform output -json | jq -r 'to_entries[] | "export TF_\(.key|ascii_upcase)=\(.value.value|@sh)"')"

# Format check in CI
terraform fmt -check -recursive -diff

# Locked provider versions
grep -A1 '^provider ' .terraform.lock.hcl

# Exit code 2 means changes are pending, 0 means none, 1 means error
terraform plan -detailed-exitcode -out=tfplan

# Plan summary counts: add, change, destroy
terraform show -json tfplan | jq '[.resource_changes[].change.actions] | flatten | group_by(.) | map({(.[0]): length}) | add'

# Attributes that force replacement, per resource
terraform show -json tfplan | jq -r '.resource_changes[] | select(.change.actions == ["delete","create"] or .change.actions == ["create","delete"]) | "\(.address): \(.change.replace_paths | map(join(".")) | join(", "))"'

# Fail CI if the plan destroys anything
terraform show -json tfplan | jq -e '[.resource_changes[] | select(.change.actions | index("delete"))] | length == 0' >/dev/null

# Drift only: which resources changed outside Terraform
terraform plan -refresh-only -detailed-exitcode -no-color | grep -E '^\s+# .* (has changed|has been deleted)'

# Every resource ID in state, with address
terraform show -json | jq -r '.values.root_module | .. | .resources? // empty | .[] | "\(.address)\t\(.values.id // "-")"'

# Providers and versions actually selected
terraform version -json | jq -r '.provider_selections | to_entries[] | "\(.key) \(.value)"'

# Modules and their sources, from the module manifest
jq -r '.Modules[] | select(.Key != "") | "\(.Key)\t\(.Source)\t\(.Version // "-")"' .terraform/modules/modules.json

# Variables without a description (awk keeps the block between variable and the closing brace)
awk '/^variable "/ {v = $2; d = 0} /^\s*description/ {d = 1} /^}/ && v != "" {if (!d) print FILENAME, v; v = ""}' *.tf

# Which state a resource with a known cloud ID lives in, across several roots
for d in envs/*; do (cd "$d" && terraform state list -id "$ID" 2>/dev/null | sed "s|^|$d: |"); done

# Move a resource into a module and verify the plan is empty
terraform state mv 'aws_instance.web' 'module.compute.aws_instance.web' && terraform plan -detailed-exitcode; echo "rc=$?"

# Untaint everything that a failed apply tainted
terraform state list | while read -r r; do terraform state show "$r" 2>/dev/null | grep -q '(tainted)' && terraform untaint "$r"; done

# Import many objects from a CSV of address,id (imperative alternative to import blocks)
while IFS=, read -r addr id; do terraform import "$addr" "$id"; done < imports.csv

# Outputs of another root as JSON, for wiring roots together in a script
terraform -chdir=envs/network output -json | jq -r '.vpc_id.value'

# Evaluate a function against the current state without a plan
terraform console <<< 'cidrsubnets("10.0.0.0/16", 4, 4, 8)'

# Validate every root module in the repository
for d in $(find . -name backend.tf -not -path '*/.terraform/*' -exec dirname {} \;); do (cd "$d" && terraform init -backend=false -input=false >/dev/null && terraform validate) || echo "FAIL $d"; done

# Run tests with a JUnit report and show only failures
terraform test -junit-xml=report.xml -json | jq -r 'select(.type == "test_run" and .test_run.status == "fail") | "\(.test_file) \(.test_run.run)"'

# Who holds the state lock (S3 lockfile backend)
aws s3 cp "s3://example-tfstate/prod/network/terraform.tfstate.tflock" - 2>/dev/null | jq .

# Age of the state snapshot and its serial
terraform state pull | jq '{serial, terraform_version, resources: (.resources | length)}'

# Sensitive outputs in plain text (prints secrets; for debugging only)
terraform output -json | jq -r 'to_entries[] | select(.value.sensitive) | .key'

# Diff two saved plans (for example before and after a provider upgrade)
diff <(terraform show -no-color before.tfplan) <(terraform show -no-color after.tfplan)

# Rebuild the lock file from scratch for all platforms
rm .terraform.lock.hcl && terraform providers lock -platform=linux_amd64 -platform=darwin_arm64

Scripts#

Plan gate for CI: runs a plan, prints a compact summary, fails when anything would be destroyed or replaced unless an allowlist file says otherwise, and leaves tfplan for the apply job.

#!/usr/bin/env bash
# usage: plan-gate.sh [allowed-destroys.txt]   run in the root module directory with credentials in the environment
set -euo pipefail

allow=${1:-/dev/null}
export TF_IN_AUTOMATION=1
terraform init -input=false >/dev/null
terraform plan -input=false -lock-timeout=5m -detailed-exitcode -out=tfplan; rc=$?
(( rc == 1 )) && { echo 'plan failed' >&2; exit 1; }
(( rc == 0 )) && { echo 'no changes'; exit 0; }

json=$(terraform show -json tfplan)
jq -r '.resource_changes[] | select(.change.actions != ["no-op"]) | "\(.change.actions | join("+"))\t\(.address)"' <<< "$json" | column -t -s $'\t'

mapfile -t destroys < <(jq -r '.resource_changes[] | select(.change.actions | index("delete")) | .address' <<< "$json")
blocked=()
for a in "${destroys[@]}"; do grep -qxF -- "$a" "$allow" || blocked+=("$a"); done
if (( ${#blocked[@]} )); then
  printf 'BLOCKED: plan destroys or replaces resources not in %s:\n' "$allow" >&2
  printf '  %s\n' "${blocked[@]}" >&2
  exit 2
fi
printf '%d changes, %d destroys (all allowed)\n' "$(jq '[.resource_changes[] | select(.change.actions != ["no-op"])] | length' <<< "$json")" "${#destroys[@]}"

Drift report across every root module under a directory: runs a refresh-only plan in each, records which resources changed outside Terraform and prints one table. Read-only apart from provider API calls; needs credentials for each root.

#!/usr/bin/env bash
# usage: drift-report.sh envs/   (each subdirectory with a backend.tf is a root module)
set -euo pipefail

base=${1:?directory of root modules}
tmp=$(mktemp -d); trap 'rm -rf -- "$tmp"' EXIT
rc=0
printf '%-20s %-8s %s\n' ROOT STATUS RESOURCES
for d in "$base"/*/; do
  [[ -f $d/backend.tf ]] || continue
  name=$(basename "$d")
  if ! terraform -chdir="$d" init -input=false >"$tmp/$name.init" 2>&1; then printf '%-20s %-8s init failed\n' "$name" ERROR; rc=1; continue; fi
  terraform -chdir="$d" plan -refresh-only -input=false -lock=false -detailed-exitcode -out="$tmp/$name.plan" >"$tmp/$name.log" 2>&1; prc=$?
  case $prc in
    0) printf '%-20s %-8s -\n' "$name" clean ;;
    2) changed=$(terraform -chdir="$d" show -json "$tmp/$name.plan" | jq -r '[.resource_drift[]? | .address] | join(", ")')
       printf '%-20s %-8s %s\n' "$name" DRIFT "$changed"; rc=1 ;;
    *) printf '%-20s %-8s see %s\n' "$name" ERROR "$tmp/$name.log"; rc=1 ;;
  esac
done
exit "$rc"

State backup before surgery: pulls the state of the current root, stores it with serial and timestamp in a backup directory, verifies the copy parses, and prints the restore command. Run it before every state mv, state rm or force-unlock.

#!/usr/bin/env bash
# usage: state-backup.sh [backup-dir]
set -euo pipefail

dir=${1:-${TF_STATE_BACKUPS:-$HOME/.terraform-state-backups}}
mkdir -p "$dir"
root=$(basename "$PWD")
ws=$(terraform workspace show)
state=$(terraform state pull)
serial=$(jq -r '.serial' <<< "$state")
lineage=$(jq -r '.lineage' <<< "$state")
[[ $serial =~ ^[0-9]+$ ]] || { echo 'state pull did not return a state file' >&2; exit 1; }
out="$dir/$root-$ws-serial$serial-$(date +%Y%m%dT%H%M%S).tfstate"
printf '%s' "$state" > "$out"
jq -e '.resources | length' "$out" >/dev/null
chmod 600 "$out"
printf 'saved %s (lineage %s, %s resources)\n' "$out" "$lineage" "$(jq '.resources | length' "$out")"
printf 'restore with: terraform state push %q\n' "$out"

Further reading#