Terraform Best Practices: Lessons from 13 Years of Infrastructure Automation
State management, module design, workspace strategy, and the anti-patterns I see repeated in teams across AWS, Azure, and GCP.
Why this keeps mattering
Terraform has been the de-facto IaC standard for years now, and yet I consistently see the same mistakes in codebases across organisations — from startups running a single AWS account to enterprises managing hundreds of accounts across multiple clouds. After 13 years of infrastructure automation work, here are the practices that have saved me (and teams I've worked with) the most pain.
1. Remote State is Non-Negotiable
Local state is fine for experimentation. The moment more than one person touches a codebase, you need remote state. Full stop.
terraform {
backend "s3" {
bucket = "my-org-tfstate"
key = "prod/networking/terraform.tfstate"
region = "eu-west-1"
encrypt = true
dynamodb_table = "terraform-locks"
}
}
The DynamoDB lock table is critical. Without it, two engineers running terraform apply simultaneously will corrupt your state file. I've seen this happen in production. It's not fun.
2. Structure Your State Thoughtfully
The single-monolithic-state-file anti-pattern is the most common mistake I see in growing teams. When everything is in one state file, a single terraform plan touches every resource — making it slow, risky, and hard to reason about.
Instead, separate state by:
- Layer — networking, compute, data, security (loosely coupled, ordered dependency)
- Environment — dev, staging, prod each get their own state
- Region — especially important for multi-region architectures
Use terraform_remote_state data sources to share outputs between layers rather than hardcoding values across state files.
3. Module Design Principles
Modules should have a single responsibility, clear input/output contracts, and sensible defaults. Avoid the "mega module" that tries to provision an entire environment in one call.
# Good: focused module
module "vpc" {
source = "./modules/vpc"
cidr_block = var.vpc_cidr
az_count = 3
}
# Anti-pattern: doing too much in one module
module "everything" {
source = "./modules/full-stack" # ❌ avoid this
}
4. Workspace Strategy
Terraform workspaces are useful but misunderstood. They are not a replacement for environment isolation — they share backend configuration and don't provide hard boundaries between environments.
My recommendation: use workspaces for lightweight environment differences (e.g. dev vs. test on the same account), and separate state backends for hard environment isolation (prod should never share a backend with dev).
5. Variable Hygiene
Never hardcode environment-specific values. Use variable files per environment:
environments/
dev.tfvars
staging.tfvars
prod.tfvars
Pass them explicitly on the CLI: terraform plan -var-file=environments/prod.tfvars. Better yet, wire this into your CI/CD pipeline so it's never a manual step.
6. Linting, Validation, and Security Scanning
Add these to every PR pipeline:
terraform fmt -check— enforces consistent formattingterraform validate— catches configuration errors before plan- tflint — catches provider-specific issues (e.g. invalid instance types)
- tfsec or Checkov — static security analysis
- Infracost — cost estimation on every PR (surprisingly valuable)
7. The Anti-Patterns That Will Hurt You
count = var.env == "prod" ? 1 : 0. It leads to index-based state keys that break on reordering.terraform state mv or terraform import, take a state backup. Always.Final thoughts
Terraform is a force multiplier for infrastructure teams, but only if it's used with discipline. The patterns above aren't clever tricks — they're the baseline that separates infrastructure that scales with your team from infrastructure that becomes a liability.
The best time to adopt these practices is when your codebase is small. The second best time is now.