KeMeT Tech
← All field notes

Terraform in Azure DevOps: Pipeline State, Drift, and What Actually Breaks

September 28, 20265 min read
terraformazure devopsiacazurepipelines

The most common failure we see on new engagements is not a broken Terraform module. It is a working module that someone has been running from their laptop for three months. The state file lives in a storage account that one engineer set up manually. The service principal is their personal one. The pipeline exists, technically, but it only runs plan; someone still does apply by hand. Then that engineer leaves.

This is not a hypothetical. We have inherited this setup on Azure DevOps engagements more times than we care to count. Here is how we rebuild it correctly.

Start With the Backend, Not the Pipeline

Almost every IaC post leads with the pipeline YAML. That is the wrong order. If your Terraform state backend is misconfigured, the pipeline is just a scheduled way to corrupt infrastructure.

For Azure, the backend is an Azure Blob container with storage account-level soft delete and versioning on. We require a dedicated storage account per environment, not per team. State for production lives in a subscription the pipeline's service principal can reach but your developers' personal accounts cannot write to.

terraform {
  required_version = ">= 1.8.0"

  required_providers {
    azurerm = {
      source  = "hashicorp/azurerm"
      version = "~> 4.0"
    }
  }

  backend "azurerm" {
    resource_group_name  = "rg-tfstate-prod"
    storage_account_name = "sttfstateprod001"
    container_name       = "tfstate"
    key                  = "platform/network.tfstate"
    use_oidc             = true
  }
}

The use_oidc = true line is the one most tutorials skip. Without it, you are passing a client secret through pipeline variables, rotating it manually, and eventually getting paged because someone's ARM_CLIENT_SECRET expired at 2 AM. Workload identity federation removes the secret entirely.

Workload Identity Federation, Not Client Secrets

Azure DevOps has supported OIDC-based service connections since mid-2023. By 2026 there is no reason to be passing ARM_CLIENT_SECRET in a variable group. The pipeline proves its identity via a short-lived JWT; Azure AD issues a token in exchange; no secret touches the YAML or the variable store.

Setting it up requires three things: a managed identity or app registration with federated credentials configured for your Azure DevOps organization and project, a service connection of type Azure Resource Manager with workload identity federation, and the corresponding environment variables set in the pipeline.

variables:
  ARM_USE_OIDC: "true"
  ARM_OIDC_TOKEN: $(System.AccessToken)
  ARM_TENANT_ID: $(AZURE_TENANT_ID)
  ARM_CLIENT_ID: $(AZURE_CLIENT_ID)
  ARM_SUBSCRIPTION_ID: $(AZURE_SUBSCRIPTION_ID)

The System.AccessToken is the Azure DevOps job token. Terraform's AzureRM provider 4.x exchanges it for an Azure AD access token automatically when ARM_USE_OIDC is set. No secret to rotate. The token is valid for the duration of the job only.

The Pipeline Structure We Actually Use

A flat pipeline with plan and apply in a single job is fine for a demo. In production you need separation: plan in one stage, human or automated gate, apply in a second stage that consumes the saved plan artifact. If apply runs with a different plan than the one reviewed, you have lost the only audit trail that matters.

stages:
  - stage: plan
    displayName: Terraform Plan
    jobs:
      - job: tf_plan
        pool:
          vmImage: ubuntu-24.04
        steps:
          - task: TerraformInstaller@1
            inputs:
              terraformVersion: "1.9.x"

          - task: TerraformTaskV4@4
            displayName: tf init
            inputs:
              provider: azurerm
              command: init
              backendServiceArm: svc-terraform-prod
              backendAzureRmResourceGroupName: rg-tfstate-prod
              backendAzureRmStorageAccountName: sttfstateprod001
              backendAzureRmContainerName: tfstate
              backendAzureRmKey: platform/network.tfstate

          - task: TerraformTaskV4@4
            displayName: tf plan
            inputs:
              provider: azurerm
              command: plan
              environmentServiceNameAzureRM: svc-terraform-prod
              commandOptions: "-out=$(Build.ArtifactStagingDirectory)/tf.plan"

          - publish: $(Build.ArtifactStagingDirectory)/tf.plan
            artifact: tfplan

  - stage: apply
    displayName: Terraform Apply
    dependsOn: plan
    condition: and(succeeded(), eq(variables['Build.SourceBranch'], 'refs/heads/main'))
    jobs:
      - deployment: tf_apply
        environment: production
        strategy:
          runOnce:
            deploy:
              steps:
                - download: current
                  artifact: tfplan

                - task: TerraformTaskV4@4
                  displayName: tf apply
                  inputs:
                    provider: azurerm
                    command: apply
                    environmentServiceNameAzureRM: svc-terraform-prod
                    commandOptions: "$(Pipeline.Workspace)/tfplan/tf.plan"

The deployment job type unlocks Azure DevOps environment approvals. Wire a required reviewer to the production environment and you get a manual gate between plan and apply with a full audit log, no custom scripting needed.

State Locking and the Blob Lease Problem

Azure Blob provides state locking through lease acquisition. Terraform holds the lease for the duration of an operation. The problem we see most: a pipeline times out or is cancelled mid-apply, the lease is not released, and the next run fails with Error: Error locking state: Error acquiring the state lock.

The fix is not to disable locking. The fix is to break the lease manually and understand why the previous run died.

# identify the lease holder
az storage blob show \
  --account-name sttfstateprod001 \
  --container-name tfstate \
  --name platform/network.tfstate \
  --query "properties.lease" \
  --output table

# break the lease
az storage blob lease break \
  --account-name sttfstateprod001 \
  --container-name tfstate \
  --blob-name platform/network.tfstate

Before breaking the lease, check the pipeline run that abandoned it. If apply was mid-flight, verify the actual Azure resource state before running again. terraform plan after a broken-lease apply will show you what diverged.

We also set the pipeline timeout to something reasonable: timeoutInMinutes: 60 on apply jobs. An uncapped job that hangs overnight holds the lease for hours.

Drift Detection as a Scheduled Pipeline

Applying on merge is table stakes. The part most teams skip is detecting drift between scheduled runs. Resources get modified in the portal, ARM deployments land outside the Terraform boundary, someone runs az cli to hotfix an outage. None of that shows up until the next apply, which either silently reverts the change or fails with a conflict.

A drift-detection pipeline runs terraform plan -detailed-exitcode on a cron, posts the diff to a Slack channel or a Teams webhook, and triggers a work item if the exit code is 2 (changes present). Exit code 0 means no drift; exit code 1 means an actual error.

schedules:
  - cron: "0 6 * * 1-5"
    displayName: Daily drift check
    branches:
      include:
        - main
    always: true

steps:
  - script: |
      terraform plan -detailed-exitcode -out=/dev/null
      EXIT=$?
      if [ $EXIT -eq 2 ]; then
        echo "##vso[task.setvariable variable=DRIFT_DETECTED]true"
      fi
    displayName: Check for drift

We have caught production config drift within hours this way on Azure Kubernetes Service node pool configurations, NSG rule additions made through the portal, and diagnostic setting removals that broke a Sentinel data connector. The Sentinel case is why this matters for security teams, not just ops.

Template Reuse Across Environments

If you have three environments and three copies of the same pipeline YAML, you will eventually fix a bug in two of them and ship the third unchanged. Azure DevOps pipeline templates solve this cleanly.

The pattern we use: a templates/ directory in a shared iac-templates repo, each template parameterized by environment, backend key, and service connection name. The environment-specific pipeline is five lines that reference the template and pass the parameters. The template itself handles init, plan, apply, and the artifact handoff.

This also means the template repo can be version-tagged. You pin to @v2.1 in environments and upgrade them independently, rather than a single change rolling to prod automatically.

See our cloud architecture and DevOps practice for the full template library and how we integrate this into trunk-based delivery workflows.

When to Call Us

If your Terraform state is still in a local backend, your pipelines run apply on every push, or you are rotating service principal secrets by hand, reach out and we will scope a one-week hardening engagement.