Segment state before anything else
A single monolithic state file across three clouds guarantees long plans, wide blast radius, and lock contention. Split state by cloud, by environment, and by lifecycle: networking changes rarely, workloads change constantly, data stores change carefully. Each boundary should own resources that are created and destroyed together.
Use remote backends with locking and versioning — S3 with DynamoDB locking, Azure Blob with lease locking, GCS with object versioning — and never let a state file live on a laptop.
Detect drift on a schedule, not during an incident
Run a scheduled plan against every state and alert on a non-empty diff. Nightly is enough for most estates; hourly is appropriate for production networking and identity. The point is to learn about drift on a Tuesday morning rather than in the middle of a change freeze.
Classify what you find. Benign drift from provider-computed fields belongs in ignore_changes or a lifecycle block. Deliberate emergency changes should be reconciled into code within days. Unexplained drift is a security signal and should be treated as one.
Adopt reality with import, not with force
When a resource exists but is unmanaged, import it and codify it. Terraform's import blocks make this reviewable in a pull request rather than a local ritual. Resist the temptation to destroy and recreate to make the plan clean — that is how drift remediation causes outages.
Version and pin providers per state. An unpinned provider upgrade will show up as drift and consume a day of investigation before anyone checks the changelog.
Enforce policy where the change happens
Run policy checks in the pipeline: no public storage buckets, mandatory tags, approved regions, size ceilings on expensive resources. Combined with restricting console write access to break-glass roles, this removes most of the drift at its source rather than cleaning it up afterward.
Key takeaways
- Split state by cloud, environment, and lifecycle to shrink blast radius.
- Schedule drift-detection plans and alert on any non-empty diff.
- Classify drift as benign, deliberate, or unexplained — treat the last as a security event.
- Use import blocks to adopt reality instead of destroying and recreating.
- Restrict console writes and enforce policy checks in the pipeline.
Need a pod that already works this way?
DevGrid Staffing assembles managed DevOps, platform, and SRE pods with the compliance and delivery practices described here built in from week one.
