Terraform
Plan and apply safely, structure modules and variables, refactor and repair state, and recover from drift, lock and provider errors in Terraform.
On this page
Cheatsheet#
| Task | Command |
|---|---|
| Initialise backend and download providers | terraform init |
| Move state to a changed backend | terraform init -migrate-state |
| Upgrade providers within constraints | terraform init -upgrade |
| Format and validate | terraform fmt -recursive && terraform validate |
| Plan to a file | terraform plan -out=tfplan |
| Apply exactly that plan | terraform apply tfplan |
| Detect drift without proposing changes | terraform plan -refresh-only |
| Recreate one resource | terraform apply -replace='aws_instance.web[0]' |
| List resources in state | terraform state list |
| Inspect one resource in state | terraform state show 'aws_instance.web[0]' |
| Back up remote state | terraform state pull > backup.tfstate |
| Release a stale lock | terraform force-unlock <lock-id> |
| Output for scripts | terraform output -raw url |
| Run module tests | terraform test |
Current release at review time: Terraform 1.16. Version-dependent features below state the version that introduced them. See the Terraform documentation.
How plan and apply work#
Terraform reads the configuration, builds a dependency graph from references between blocks, refreshes state by asking providers for the current attributes of every tracked resource, then diffs desired against current and proposes create, update, replace or destroy actions. Independent resources are applied in parallel (10 at a time by default, -parallelism=n).
State maps each resource address (aws_instance.web[0]) to a real object ID. It is the only link between configuration and the real world: a resource missing from state is invisible to Terraform, and a resource removed from configuration but still in state is scheduled for destruction.
Three things can disagree: configuration, state and reality.
| Command | Compares | Changes infrastructure |
|---|---|---|
terraform plan | Configuration against refreshed state | No |
terraform plan -refresh-only | State against reality (drift) | No |
terraform apply -refresh-only | Writes reality into state | No, state only |
terraform apply | Configuration against refreshed state | Yes |
terraform plan -out=tfplan
terraform show tfplan # human-readable review of the saved plan
terraform apply tfplan # applies exactly what was reviewed, or fails if state moved onAlways use plan -out then apply <file> in automation. apply without a plan file re-plans, so what runs can differ from what was reviewed. A saved plan fails to apply if the state changed after it was created.
Plan symbols: + create, - destroy, ~ update in place, -/+ destroy then create, +/- create then destroy (create_before_destroy), <= read a data source. Look for # forces replacement next to an attribute to see why a resource is being replaced.
Configuration and providers#
terraform {
required_version = ">= 1.11"
required_providers {
aws = { source = "hashicorp/aws", version = "~> 6.0" }
}
backend "s3" {
bucket = "example-tfstate"
key = "prod/network/terraform.tfstate"
region = "ap-southeast-2"
use_lockfile = true # S3-native locking; DynamoDB locking is deprecated
encrypt = true
}
}
provider "aws" {
region = "ap-southeast-2"
}
provider "aws" {
alias = "us_east_1" # for resources that must live in us-east-1, such as CloudFront certificates
region = "us-east-1"
}terraform init writes .terraform.lock.hcl with the exact provider versions and checksums selected. Commit it: it is what makes two machines use the same provider build. init -upgrade moves to the newest version the constraints allow and rewrites the lock file.
The dynamodb_table backend argument still works but is deprecated in favour of use_lockfile; both can be set during migration. See the S3 backend.
Resources and meta-arguments#
resource "aws_instance" "web" {
for_each = var.web_nodes # map of name => { subnet_id = ... }
ami = data.aws_ami.al2023.id
instance_type = var.instance_type
subnet_id = each.value.subnet_id
vpc_security_group_ids = [aws_security_group.web.id]
tags = merge(var.tags, { Name = "${var.name}-${each.key}" })
lifecycle {
create_before_destroy = true
ignore_changes = [ami] # an image pipeline owns this attribute
precondition {
condition = var.instance_type != "t2.micro"
error_message = "t2.micro is not permitted in production."
}
}
}| Meta-argument | Effect |
|---|---|
count | Instances indexed by position (web[0]); removing a middle element shifts every later index |
for_each | Instances keyed by map key or set value (web["a"]); adding or removing a key touches only that key |
depends_on | Explicit ordering when no expression reference exists |
provider | Selects an aliased provider configuration (another region or account) |
lifecycle.create_before_destroy | Creates the replacement before destroying the old object |
lifecycle.prevent_destroy | Fails any plan that would destroy the resource; use on databases and state buckets |
lifecycle.ignore_changes | Stops Terraform reverting changes to listed attributes |
lifecycle.replace_triggered_by | Replaces this resource when a referenced resource or attribute changes |
Prefer for_each for anything with a natural key. With count, deleting the first list element destroys and recreates every resource after it. for_each keys must be known at plan time, so they cannot come from attributes of resources that do not exist yet.
Variables, outputs and precedence#
variable "instance_type" {
type = string
default = "t3.small"
description = "EC2 instance size for the web tier"
validation {
condition = can(regex("^t3\\.", var.instance_type))
error_message = "Only t3 sizes are approved here."
}
}
output "url" {
value = "https://${aws_lb.this.dns_name}"
description = "Public endpoint"
}Precedence, lowest to highest: the variable’s default, TF_VAR_<name> environment variables, terraform.tfvars, terraform.tfvars.json, *.auto.tfvars and *.auto.tfvars.json in lexical order, then -var and -var-file in command-line order. A later source replaces an earlier value entirely; maps are not merged.
Keeping secrets out of state#
sensitive = true hides a value from CLI output only. It is still written in plain text to state and saved plan files.
| Feature | Version | Stored in state or plan |
|---|---|---|
sensitive = true on variables and outputs | 0.15 | Yes, redacted in output only |
ephemeral = true on variables and outputs, ephemeral resource blocks | 1.10 | No |
Write-only resource arguments (usually named *_wo, paired with a *_wo_version) | 1.11 | No |
ephemeral "random_password" "db" {
length = 32
}
resource "aws_db_instance" "main" {
# ...
password_wo = ephemeral.random_password.db.result
password_wo_version = 1 # bump to push a new password
}Write-only arguments exist only where the provider implements them; check the resource documentation. See managing sensitive data and, for issuing secrets at apply time, Vault.
State is a secret
Encrypt the backend, restrict read access to the people and pipelines that run Terraform, turn on bucket versioning so a bad write can be rolled back, and never commit terraform.tfstate or *.tfvars files holding credentials.
Modules#
module "network" {
source = "git::https://github.com/example/tf-modules.git//network?ref=v1.4.0"
name = "prod"
cidr = "10.20.0.0/16"
az_count = 3
}
module "vpc" {
source = "terraform-aws-modules/vpc/aws" # registry module
version = "~> 6.0"
# ...
}Pin module sources to a tag or commit (?ref=) or a registry version. A floating branch lets a plan change because someone else merged. Run terraform init (or init -upgrade) after changing a source or version.
Keep modules to inputs, resources and outputs. A module that hard-codes naming, tagging or environment decisions is harder to reuse than one that accepts them as variables. terraform test (1.6+) runs *.tftest.hcl files that plan or apply a module and assert on the result.
Refactoring and repairing state#
Prefer configuration blocks to state commands: they appear in the plan, are reviewed in a pull request and apply the same way in every workspace.
| Goal | Declarative (reviewed in plan) | Imperative (immediate, unreviewed) |
|---|---|---|
| Rename or move into a module | moved block (1.1+) | terraform state mv |
| Stop managing, keep the object | removed block with lifecycle { destroy = false } (1.7+) | terraform state rm |
| Adopt an existing object | import block (1.5+; for_each 1.7+) | terraform import |
moved {
from = aws_instance.web
to = module.compute.aws_instance.web
}
removed {
from = aws_instance.legacy
lifecycle {
destroy = false
}
}
import {
to = aws_s3_bucket.logs
id = "example-logs-bucket"
}terraform plan -generate-config-out=generated.tf # writes resource blocks for import targets that lack oneReview generated configuration before applying; it contains every attribute, including defaults you should delete.
Back up before state surgery
Run terraform state pull > "state-$(date +%Y%m%dT%H%M%S).json" before any state mv, state rm, state push or force-unlock. These commands change remote state immediately with no plan.
terraform state list
terraform state show 'aws_instance.web["a"]'
terraform state mv 'aws_instance.web' 'aws_instance.api'
terraform state rm 'aws_instance.legacy' # forget it; the object keeps running
terraform import 'aws_instance.web["a"]' i-0abc123Quote addresses that contain [ or " so the shell passes them unchanged.
moved, import and removed in depth#
moved blocks chain: when a resource is renamed twice across releases, keep both blocks so any state at either old address migrates. They work across module boundaries and for whole modules, and for count to for_each conversions where each index needs its own block.
moved { # whole module rename; every resource inside moves with it
from = module.net
to = module.network
}
moved { # count to for_each: one block per index
from = aws_subnet.private[0]
to = aws_subnet.private["a"]
}
moved {
from = aws_subnet.private[1]
to = aws_subnet.private["b"]
}
import { # bulk import driven by a map (1.7+)
for_each = var.existing_buckets # { logs = "example-logs", assets = "example-assets" }
to = aws_s3_bucket.managed[each.key]
id = each.value
}
import {
to = aws_route53_record.www
id = "Z0123456789ABC_www.example.com_A" # ID formats are provider-specific; check the resource docs "Import" section
} # 1.12+: providers may accept an identity {...} object instead of an id string
removed {
from = module.legacy # also works for whole modules
lifecycle { destroy = false }
}A plan shows # (moved from ...) and # (imported from ...) lines, and an import for an object whose attributes differ from the configuration proposes an update in the same plan; read that update before applying. Once applied, delete the import and moved blocks in a later change (they are harmless while present, but moved blocks whose from no longer exists in anyone’s state are noise). terraform plan -generate-config-out refuses to overwrite an existing file and only generates for import blocks whose target has no configuration.
State surgery#
State is JSON with a serial and lineage. The backend rejects a push whose serial is not newer or whose lineage differs, which is the safety net when hand-editing. Every operation below rewrites remote state immediately.
terraform state pull > state.json # always first; keep this copy until the next successful plan
terraform state list -id i-0abc123 # which address holds an object with that ID
terraform state show -no-color 'module.db.aws_db_instance.main' | grep -E '^\s+(arn|id) '
terraform state mv 'module.old.aws_instance.web' 'module.new.aws_instance.web'
terraform state rm 'module.legacy' # forgets every resource under the module
terraform state replace-provider registry.terraform.io/-/aws registry.terraform.io/hashicorp/aws # after the 0.13 provider namespace change or a fork
terraform apply -replace='aws_instance.web["a"]' # supersedes terraform taint
terraform untaint 'aws_instance.web["a"]' # clear a taint left by a failed create
terraform apply -refresh-only # accept reality into state without touching configuration
terraform state push state.json # upload an edited copy; refuses older serial or different lineage
terraform state push -force state.json # overrides both checks; last resortMoving a resource between two root modules (two backends) has no single command. Pull both states, move with the local-file form, then push both:
cd ../network && terraform state pull > /tmp/net.json && cd ../compute && terraform state pull > /tmp/comp.json
terraform state mv -state=/tmp/net.json -state-out=/tmp/comp.json 'aws_security_group.web' 'aws_security_group.web'
cd ../network && terraform state push /tmp/net.json && cd ../compute && terraform state push /tmp/comp.json
# then add the resource block to compute, delete it from network, and plan both: each should show no changesEditing the JSON by hand is occasionally necessary (a provider bug wrote an impossible attribute, or a resource must be nudged to a new schema version). Bump serial, keep lineage, and validate with terraform show state.json before pushing. A lost state file is recovered by importing every object, which is why terraform state list output belongs in the run logs of any pipeline.
Workspaces vs directories#
Two ways to run one configuration against several environments:
| CLI workspaces | One directory per environment | |
|---|---|---|
| State | Same backend, separate state per workspace (env:/<name>/ prefix on S3) | Separate backend configuration and state per directory |
| Credentials | Shared; the same principal can touch every environment | Can differ per directory (separate roles, accounts, subscriptions) |
| Configuration drift between environments | Impossible: one set of files | Possible; controlled by sharing modules and diffing tfvars |
Blast radius of a wrong apply | High: terraform workspace select is easy to forget | Low: -chdir=envs/prod is explicit |
| Suits | Ephemeral copies (review environments, per-branch stacks) | Long-lived environments with different sizes, providers and approvers |
envs/
prod/ main.tf backend.tf prod.tfvars # root module: a handful of module calls and provider config
staging/ main.tf backend.tf staging.tfvars
modules/
network/ compute/ database/ # all environment differences arrive through variablesterraform -chdir=envs/prod init -backend-config=prod.s3.tfbackend # partial backend config kept out of source when it holds account-specific values
terraform -chdir=envs/prod plan -var-file=prod.tfvars -out=tfplan
terraform workspace new pr-1234 && terraform apply -var-file=review.tfvars -auto-approve # ephemeral copy
terraform workspace select default && terraform workspace delete pr-1234 # refuses while its state is non-empty; destroy firstterraform.workspace in expressions (name = "app-${terraform.workspace}") makes workspaces usable but also hides the environment from the reader of the configuration. If a configuration needs count = terraform.workspace == "prod" ? 3 : 1, it wants directories.
Tests#
terraform test runs *.tftest.hcl files in the root or tests/ directory. Each run block plans (command = plan) or applies (command = apply, the default) the module with the given variables and checks assert conditions; applied resources are destroyed when the file finishes. Mock providers (1.7+) return fabricated values so unit tests need no credentials.
# tests/network.tftest.hcl
variables {
name = "test"
cidr = "10.99.0.0/16"
}
mock_provider "aws" {} # every resource and data source returns generated values
run "creates_one_subnet_per_az" {
command = plan
variables { az_count = 3 }
assert {
condition = length(aws_subnet.private) == 3
error_message = "expected 3 private subnets, got ${length(aws_subnet.private)}"
}
}
run "rejects_small_cidr" {
command = plan
variables { cidr = "10.99.0.0/28" }
expect_failures = [var.cidr] # the variable's validation block must reject it
}
run "outputs_match" {
command = plan
assert {
condition = output.vpc_cidr == var.cidr
error_message = "output must echo the input CIDR"
}
}terraform test # every test file
terraform test -filter=tests/network.tftest.hcl # one file
terraform test -verbose # print the plan or state for each run
terraform test -junit-xml=report.xml # JUnit report for CIUse command = plan with mocks for fast unit checks of logic (counts, names, validation), and a small number of real apply runs against a sandbox account for integration. A run block can reference outputs of an earlier run (run.setup.some_output) and can load a helper module with module { source = "./tests/setup" } to create prerequisites. override_resource and override_data blocks replace one object’s attributes without mocking the whole provider.
Provider lock and installation#
.terraform.lock.hcl records, per provider, the selected version, the constraints in effect and checksums (h1: for the zip on this platform, zh: for every platform’s zip from the registry). terraform init on a platform whose h1: hash is missing fails with a checksum error unless the lock file was generated for that platform too.
terraform providers # required providers and which module requires them
terraform providers lock -platform=linux_amd64 -platform=darwin_arm64 -platform=linux_arm64 # record hashes for every platform CI and laptops use
terraform init -upgrade # newest versions within constraints; rewrites the lock file
terraform providers mirror ./mirror # download providers for an air-gapped or rate-limited environment
terraform providers schema -json | jq '.provider_schemas | keys' # inspect installed provider schemas# ~/.terraformrc: reuse downloaded providers across working directories and prefer a local mirror
plugin_cache_dir = "$HOME/.terraform.d/plugin-cache"
provider_installation {
filesystem_mirror { path = "/opt/terraform/mirror", include = ["registry.terraform.io/hashicorp/*"] }
direct { exclude = ["registry.terraform.io/hashicorp/*"] }
}TF_PLUGIN_CACHE_DIR does the same as plugin_cache_dir from the environment. The cache is not safe for concurrent init runs of different working directories on the same host (a known limitation); serialise them in CI or use a mirror. Version constraints belong in required_providers of the root module; child modules should state only the minimum they need (>= 5.0) so the root decides.
Expressions worth knowing#
locals {
env_tags = merge(var.tags, { Environment = var.env, ManagedBy = "terraform" })
subnets = { for az in var.azs : az => cidrsubnet(var.cidr, 8, index(var.azs, az)) } # for expression over a list into a map
public = [for s in aws_subnet.all : s.id if s.tags["tier"] == "public"] # filter
cfg = yamldecode(file("${path.module}/config.yaml"))
name = coalesce(var.name_override, "${var.project}-${var.env}")
}
dynamic "ingress" { # repeat a nested block per element
for_each = var.ingress_rules
content {
from_port = ingress.value.port
to_port = ingress.value.port
protocol = "tcp"
cidr_blocks = ingress.value.cidrs
}
}
check "endpoint_answers" { # post-apply assertion that does not block the apply (1.5+)
data "http" "health" { url = "https://${aws_lb.this.dns_name}/health" }
assert {
condition = data.http.health.status_code == 200
error_message = "health endpoint returned ${data.http.health.status_code}"
}
}terraform console <<< 'cidrsubnet("10.0.0.0/16", 8, 3)' # evaluate expressions against the current state
terraform console <<< 'keys(module.network.subnet_ids)'try(expr, fallback) and can(expr) handle optional attributes; one() turns a zero-or-one element list into a value or null; sensitive() and nonsensitive() adjust redaction. templatefile() renders a file with variables and is preferable to long heredocs for user data and policies.
Workspaces and environment layout#
CLI workspaces keep separate state files under one backend configuration (for S3, under the env:/ key prefix). They suit short-lived copies of identical infrastructure, such as a feature branch stack. Production and development usually differ in size, providers and credentials, so give them separate root directories with separate backends and credentials instead.
terraform workspace list
terraform workspace new feature-x
terraform workspace select defaultterraform.workspace holds the current name for use in expressions.
Troubleshooting#
| Symptom | Cause | Fix |
|---|---|---|
Error acquiring the state lock | Another run is active, or a crashed run left the lock | Confirm no run is active (CI, colleagues), then terraform force-unlock <id> |
Saved plan is stale | State changed after the plan was written | Plan again |
| Plan replaces many resources after a list edit | count index shift | Switch to for_each and add moved blocks |
| Plan replaces resources after a provider upgrade | Changed defaults or schema in the new major version | Read the provider upgrade guide; pin the old version until handled |
| Resource exists but plan wants to create it | Not in state (created by hand, or state lost) | Import it |
| Resource gone but plan wants to update it | Deleted outside Terraform | terraform apply -refresh-only, then plan |
| Perpetual diff on every plan | API normalises the value (case, JSON ordering, defaults) | Match the normalised form, or ignore_changes |
Provider produced inconsistent result after apply | Provider bug or eventually consistent API | Re-run; report upstream if it persists |
Invalid for_each argument ... will be known only after apply | Keys depend on unknown values | Build keys from variables or static values |
Cycle: error | Two resources reference each other, often security groups | Split rules into separate rule resources |
Inconsistent dependency lock file | Provider added without re-running init | terraform init (or init -upgrade) and commit the lock file |
| Slow plans | Thousands of resources in one state; refresh calls every API | Split state by lifecycle and blast radius; -refresh=false for a quick look only |
Failed to install provider ... checksum mismatch or doesn't match any of the checksums | Lock file has hashes for another platform only | terraform providers lock -platform=... for every platform, commit the lock file |
Backend configuration changed | Backend block differs from .terraform/terraform.tfstate | terraform init -migrate-state to move, or -reconfigure to point elsewhere without copying |
import block plans an update as well as the import | Configuration differs from the real object’s attributes | Align the configuration with terraform state show output after import, or accept the change knowingly |
Moved object still exists at ... from address | Both from and to addresses exist in state | Remove one with state rm after confirming which is real |
terraform test fails with credentials errors | run blocks default to command = apply against real providers | command = plan plus mock_provider, or supply sandbox credentials |
workspace delete refuses | Workspace state is not empty | terraform destroy in that workspace first, or -force to abandon the objects (they keep running) |
Error: Unsupported attribute after a provider upgrade | Attribute renamed or removed in the new major version | Provider changelog and upgrade guide; terraform providers schema -json to see the current schema |
| Plan wants to destroy everything | Wrong workspace, wrong backend key, or empty state after a failed migration | terraform workspace show, terraform state list, stop and compare with the backup |
Error: Invalid function argument on file() | Path relative to the working directory, not the module | Use ${path.module}/file |
ephemeral value used in a non-ephemeral context | Ephemeral values may only feed write-only arguments, provider config or other ephemerals | Restructure; do not copy the value into a normal attribute |
TF_LOG=DEBUG terraform plan 2>debug.log # core and provider logs; may contain secrets
TF_LOG_PROVIDER=TRACE terraform apply # provider API calls only
terraform providers # provider requirements per module
terraform graph | dot -Tsvg > graph.svg # needs graphviz-target is for recovery
-target plans a subset of the graph and can leave state inconsistent with configuration. Use it to get out of a broken state, then run a full plan.
Oneliners#
# Every change in a saved plan, one line each
terraform show -json tfplan | jq -r '.resource_changes[] | select(.change.actions != ["no-op"]) | "\(.change.actions|join(","))\t\(.address)"'
# Only the deletes and replacements: the review that matters
terraform show -json tfplan | jq -r '.resource_changes[] | select(.change.actions | index("delete")) | .address'
# Resource counts by type
terraform state list | sed 's/\[.*//' | awk -F. '{print $(NF-1)}' | sort | uniq -c | sort -rn
# Outputs into environment variables
eval "$(terraform output -json | jq -r 'to_entries[] | "export TF_\(.key|ascii_upcase)=\(.value.value|@sh)"')"
# Format check in CI
terraform fmt -check -recursive -diff
# Locked provider versions
grep -A1 '^provider ' .terraform.lock.hcl
# Exit code 2 means changes are pending, 0 means none, 1 means error
terraform plan -detailed-exitcode -out=tfplan
# Plan summary counts: add, change, destroy
terraform show -json tfplan | jq '[.resource_changes[].change.actions] | flatten | group_by(.) | map({(.[0]): length}) | add'
# Attributes that force replacement, per resource
terraform show -json tfplan | jq -r '.resource_changes[] | select(.change.actions == ["delete","create"] or .change.actions == ["create","delete"]) | "\(.address): \(.change.replace_paths | map(join(".")) | join(", "))"'
# Fail CI if the plan destroys anything
terraform show -json tfplan | jq -e '[.resource_changes[] | select(.change.actions | index("delete"))] | length == 0' >/dev/null
# Drift only: which resources changed outside Terraform
terraform plan -refresh-only -detailed-exitcode -no-color | grep -E '^\s+# .* (has changed|has been deleted)'
# Every resource ID in state, with address
terraform show -json | jq -r '.values.root_module | .. | .resources? // empty | .[] | "\(.address)\t\(.values.id // "-")"'
# Providers and versions actually selected
terraform version -json | jq -r '.provider_selections | to_entries[] | "\(.key) \(.value)"'
# Modules and their sources, from the module manifest
jq -r '.Modules[] | select(.Key != "") | "\(.Key)\t\(.Source)\t\(.Version // "-")"' .terraform/modules/modules.json
# Variables without a description (awk keeps the block between variable and the closing brace)
awk '/^variable "/ {v = $2; d = 0} /^\s*description/ {d = 1} /^}/ && v != "" {if (!d) print FILENAME, v; v = ""}' *.tf
# Which state a resource with a known cloud ID lives in, across several roots
for d in envs/*; do (cd "$d" && terraform state list -id "$ID" 2>/dev/null | sed "s|^|$d: |"); done
# Move a resource into a module and verify the plan is empty
terraform state mv 'aws_instance.web' 'module.compute.aws_instance.web' && terraform plan -detailed-exitcode; echo "rc=$?"
# Untaint everything that a failed apply tainted
terraform state list | while read -r r; do terraform state show "$r" 2>/dev/null | grep -q '(tainted)' && terraform untaint "$r"; done
# Import many objects from a CSV of address,id (imperative alternative to import blocks)
while IFS=, read -r addr id; do terraform import "$addr" "$id"; done < imports.csv
# Outputs of another root as JSON, for wiring roots together in a script
terraform -chdir=envs/network output -json | jq -r '.vpc_id.value'
# Evaluate a function against the current state without a plan
terraform console <<< 'cidrsubnets("10.0.0.0/16", 4, 4, 8)'
# Validate every root module in the repository
for d in $(find . -name backend.tf -not -path '*/.terraform/*' -exec dirname {} \;); do (cd "$d" && terraform init -backend=false -input=false >/dev/null && terraform validate) || echo "FAIL $d"; done
# Run tests with a JUnit report and show only failures
terraform test -junit-xml=report.xml -json | jq -r 'select(.type == "test_run" and .test_run.status == "fail") | "\(.test_file) \(.test_run.run)"'
# Who holds the state lock (S3 lockfile backend)
aws s3 cp "s3://example-tfstate/prod/network/terraform.tfstate.tflock" - 2>/dev/null | jq .
# Age of the state snapshot and its serial
terraform state pull | jq '{serial, terraform_version, resources: (.resources | length)}'
# Sensitive outputs in plain text (prints secrets; for debugging only)
terraform output -json | jq -r 'to_entries[] | select(.value.sensitive) | .key'
# Diff two saved plans (for example before and after a provider upgrade)
diff <(terraform show -no-color before.tfplan) <(terraform show -no-color after.tfplan)
# Rebuild the lock file from scratch for all platforms
rm .terraform.lock.hcl && terraform providers lock -platform=linux_amd64 -platform=darwin_arm64Scripts#
Plan gate for CI: runs a plan, prints a compact summary, fails when anything would be destroyed or replaced unless an allowlist file says otherwise, and leaves tfplan for the apply job.
#!/usr/bin/env bash
# usage: plan-gate.sh [allowed-destroys.txt] run in the root module directory with credentials in the environment
set -euo pipefail
allow=${1:-/dev/null}
export TF_IN_AUTOMATION=1
terraform init -input=false >/dev/null
terraform plan -input=false -lock-timeout=5m -detailed-exitcode -out=tfplan; rc=$?
(( rc == 1 )) && { echo 'plan failed' >&2; exit 1; }
(( rc == 0 )) && { echo 'no changes'; exit 0; }
json=$(terraform show -json tfplan)
jq -r '.resource_changes[] | select(.change.actions != ["no-op"]) | "\(.change.actions | join("+"))\t\(.address)"' <<< "$json" | column -t -s $'\t'
mapfile -t destroys < <(jq -r '.resource_changes[] | select(.change.actions | index("delete")) | .address' <<< "$json")
blocked=()
for a in "${destroys[@]}"; do grep -qxF -- "$a" "$allow" || blocked+=("$a"); done
if (( ${#blocked[@]} )); then
printf 'BLOCKED: plan destroys or replaces resources not in %s:\n' "$allow" >&2
printf ' %s\n' "${blocked[@]}" >&2
exit 2
fi
printf '%d changes, %d destroys (all allowed)\n' "$(jq '[.resource_changes[] | select(.change.actions != ["no-op"])] | length' <<< "$json")" "${#destroys[@]}"Drift report across every root module under a directory: runs a refresh-only plan in each, records which resources changed outside Terraform and prints one table. Read-only apart from provider API calls; needs credentials for each root.
#!/usr/bin/env bash
# usage: drift-report.sh envs/ (each subdirectory with a backend.tf is a root module)
set -euo pipefail
base=${1:?directory of root modules}
tmp=$(mktemp -d); trap 'rm -rf -- "$tmp"' EXIT
rc=0
printf '%-20s %-8s %s\n' ROOT STATUS RESOURCES
for d in "$base"/*/; do
[[ -f $d/backend.tf ]] || continue
name=$(basename "$d")
if ! terraform -chdir="$d" init -input=false >"$tmp/$name.init" 2>&1; then printf '%-20s %-8s init failed\n' "$name" ERROR; rc=1; continue; fi
terraform -chdir="$d" plan -refresh-only -input=false -lock=false -detailed-exitcode -out="$tmp/$name.plan" >"$tmp/$name.log" 2>&1; prc=$?
case $prc in
0) printf '%-20s %-8s -\n' "$name" clean ;;
2) changed=$(terraform -chdir="$d" show -json "$tmp/$name.plan" | jq -r '[.resource_drift[]? | .address] | join(", ")')
printf '%-20s %-8s %s\n' "$name" DRIFT "$changed"; rc=1 ;;
*) printf '%-20s %-8s see %s\n' "$name" ERROR "$tmp/$name.log"; rc=1 ;;
esac
done
exit "$rc"State backup before surgery: pulls the state of the current root, stores it with serial and timestamp in a backup directory, verifies the copy parses, and prints the restore command. Run it before every state mv, state rm or force-unlock.
#!/usr/bin/env bash
# usage: state-backup.sh [backup-dir]
set -euo pipefail
dir=${1:-${TF_STATE_BACKUPS:-$HOME/.terraform-state-backups}}
mkdir -p "$dir"
root=$(basename "$PWD")
ws=$(terraform workspace show)
state=$(terraform state pull)
serial=$(jq -r '.serial' <<< "$state")
lineage=$(jq -r '.lineage' <<< "$state")
[[ $serial =~ ^[0-9]+$ ]] || { echo 'state pull did not return a state file' >&2; exit 1; }
out="$dir/$root-$ws-serial$serial-$(date +%Y%m%dT%H%M%S).tfstate"
printf '%s' "$state" > "$out"
jq -e '.resources | length' "$out" >/dev/null
chmod 600 "$out"
printf 'saved %s (lineage %s, %s resources)\n' "$out" "$lineage" "$(jq '.resources | length' "$out")"
printf 'restore with: terraform state push %q\n' "$out"Further reading#
- Terraform CLI commands: every subcommand and flag, including
state,providers lock,testandimport. - Refactoring with moved blocks and import blocks: supported address forms and generation of configuration.
- Dependency lock file: what the hashes mean and how multi-platform locking works.
- Tests:
runblocks, mocks, overrides andexpect_failures. - Workspaces: when they fit and how backends store them.
- Manage sensitive data:
sensitive,ephemeraland write-only arguments.