KeMeT Tech
← All field notes

A Terraform Azure Example That Covers State, Providers, and Drift

October 3, 20265 min read
terraformazureiaccloud-architecturedevops

Most people searching "terraform azure example" already have a resource "azurerm_resource_group" block somewhere. What they're actually missing is everything around it: which provider to use, where state lives, and what happens six months later when a teammate changes a storage account setting in the portal and terraform plan suddenly wants to destroy it. We've rebuilt enough client Terraform repos from scratch to know the example that matters isn't the resource block. It's the scaffolding.

Pick the provider before you pick the resource

Azure has four Terraform providers worth knowing, and picking the wrong one early costs a rewrite later. Microsoft's overview lays them out:

  • AzureRM: stable Azure resources, virtual machines, storage accounts, networking interfaces. This is the default for almost everything.
  • AzAPI: talks to Azure Resource Manager APIs directly, so you get access to Azure's newest capabilities without waiting on a provider update.
  • AzureAD: Microsoft Entra resources, groups, users, service principals, applications.
  • AzureDevops: Azure DevOps agents, repositories, projects, pipelines, queries.

The AzureRM-vs-AzAPI question comes up constantly on client engagements, enough that Microsoft and HashiCorp published a joint statement on when to reach for each. Our rule of thumb: start in AzureRM. Drop to AzAPI only for the specific resource or property that AzureRM hasn't caught up on yet, not as a wholesale replacement. Mixing the two across an entire module makes state and plans harder to reason about than they need to be.

A working example: remote state first

Every "getting started" example that stores state as a local .tfstate file is a trap. Local state doesn't work in a team setting, it can carry secrets in plain text, and it's one rm -rf away from gone. Microsoft's guide to storing Terraform state in Azure Storage gives the full bootstrap, and this is the example worth actually running before anything else:

RESOURCE_GROUP_NAME=tfstate
STORAGE_ACCOUNT_NAME=tfstate$RANDOM
CONTAINER_NAME=tfstate

az group create --name $RESOURCE_GROUP_NAME --location eastus
az storage account create --resource-group $RESOURCE_GROUP_NAME --name $STORAGE_ACCOUNT_NAME --sku Standard_LRS --encryption-services blob
az storage container create --name $CONTAINER_NAME --account-name $STORAGE_ACCOUNT_NAME

Or do it in Terraform itself, which is the version we actually keep in client repos:

terraform {
  required_providers {
    azurerm = {
      source  = "hashicorp/azurerm"
      version = "~>3.0"
    }
  }
}

provider "azurerm" {
  features {}
}

resource "random_string" "resource_code" {
  length  = 5
  special = false
  upper   = false
}

resource "azurerm_resource_group" "tfstate" {
  name     = "tfstate"
  location = "East US"
}

resource "azurerm_storage_account" "tfstate" {
  name                              = "tfstate${random_string.resource_code.result}"
  resource_group_name               = azurerm_resource_group.tfstate.name
  location                          = azurerm_resource_group.tfstate.location
  account_tier                      = "Standard"
  account_replication_type          = "LRS"
  allow_nested_items_to_be_public   = false
}

resource "azurerm_storage_container" "tfstate" {
  name                  = "tfstate"
  storage_account_id    = azurerm_storage_account.tfstate.id
  container_access_type = "private"
}

Save that as create-remote-storage.tf, run terraform init then terraform apply, and you have a backend before you have a single application resource. In this bootstrap example, authentication to the storage account uses an access key; Microsoft's own guidance is to evaluate the available azurerm backend auth options and pick the most secure one for your actual production use, not just copy the quickstart verbatim.

The part everyone skips: drift on the state bucket itself

Here's the failure mode we've walked into on client engagements: someone converts the storage account's redundancy setting in the portal, outside Terraform, and the next plan wants to delete and recreate the account that holds your state. Microsoft documents this directly as a known risk with storage account conversions. Their mitigation, which we now apply as standing practice on every backend module:

  • Turn off automatic approval for Terraform deployments during a conversion window and actually read every plan for unexpected replacement.
  • Set prevent_destroy = true in the storage account's lifecycle block.
  • Temporarily set ignore_changes on account_replication_type while the conversion is in flight.
  • After the conversion finishes: run terraform apply -refresh-only, update the config to match the new values, remove the ignore_changes entry, then run plan again and confirm it reports nothing unexpected.

That sequence is tedious enough that most teams skip it until the first outage teaches them otherwise. Bake it into your runbook now instead of after.

What you actually wire up after the backend

Once state is solid, the Terraform-on-Azure docs index a long list of quickstarts worth knowing exist even if you don't need all of them today: AKS clusters, Linux and Windows VMs, Key Vault, Application Gateway, Azure SQL Database, API Management, Front Door Standard/Premium, Container Instances, and on the networking side, IP Groups, Availability Zones, NAT Gateway, and private endpoints. If you're starting from an existing Azure footprint rather than greenfield, Azure Export for Terraform will generate HCL from resources that already exist, which is faster than hand-transcribing a portal deployment into .tf files. We lean on this a lot on migration engagements where infrastructure predates any IaC discipline.

Pin your provider version, because 2.0 actually broke things

This isn't theoretical. When HashiCorp and Microsoft shipped the AzureRM 2.0 provider, it was a genuine breaking release: deprecated fields were removed outright, and the provider started requiring that pre-existing resources be imported into state before Terraform would manage them, instead of silently re-creating them. It also introduced resource-level custom timeouts:

resource "azurerm_resource_group" "example" {
  name     = "example-resource-group"
  location = "West Europe"

  timeouts {
    create = "10m"
    delete = "30m"
  }
}

and split the old catch-all VM resources into OS-specific ones (azurerm_linux_virtual_machine, azurerm_windows_virtual_machine, and their scale-set equivalents). None of that mattered until someone ran terraform apply on an unpinned provider in CI and got a different plan than they expected. Pin the version in required_providers the way the state-storage example above does it, and treat a major version bump as a change you review, not one you absorb automatically.

Next steps

If you're standing up Terraform on Azure for the first time, or untangling a repo where state, provider versions, and drift have already gotten away from you, our cloud architecture practice does this as day-to-day work, and you can reach us through /contact.