
Last night at 20:30 UTC, Microsoft's own maintenance broke its own cloud. An infrastructure OS servicing activity — the kind of background patching every ops team does — took down or degraded ExpressRoute gateways, VPN gateways, and Azure VMware services across 18 regions. By 23:13 UTC the ExpressRoute gateways were recovering; Microsoft called it mitigated at 03:30 UTC this morning and promised a deeper explanation later.
If you run hybrid connectivity, this is the kind of event that pages you at dinner and makes your ISP the first suspect. The durable lesson isn't "Azure broke." It's that you can't patch the provider, can't roll back the provider's change, and can't even watch them do it. The only part of a provider outage you control is your own runbook for telling a provider problem apart from yours — quickly, with evidence, before the war room forms.
What actually happened
Microsoft's status page put the onset at 20:30 UTC on September 30, 2026. Customers saw degraded or interrupted connectivity on ExpressRoute gateways, VPN gateways, and Azure VMware Solution, plus failures and delays in network management operations. The Register counted 18 impacted regions — West US, West US 3, North Europe, West Europe, France Central, UK West, UK South, Switzerland North, Southeast Asia, East Asia, Japan West, Korea Central, South Africa North, UAE North, Mexico Central, Germany North, South India, and Jio India Central. Azure VMware services were unaffected in the last four of those.
Two details matter for how you think about this. First, some VPN gateways lost redundancy rather than connectivity outright — reduced to a single instance, one failure away from down, with nothing visibly broken from the user's side. Second, management-plane components kept failing after the data plane recovered: resource creation, updates, and rule changes on things like Azure Firewall and Application Gateway were failing or delayed in a subset of regions. Microsoft said its engineers were restoring those components from "alternate healthy instances." That is ops language for "we failed over to the spare parts and we're watching."
The cause, as Microsoft describes it, is a correlation between the event and its own infrastructure OS servicing activity. Redmond paused that servicing to stop the bleeding. A root cause has not been published as of this morning. "Correlated with our own patching" is about as self-inflicted as a cloud outage gets.
The two traps this kind of outage sets
Trap one: you troubleshoot your own side first. The symptoms — tunnels flapping, ExpressRoute prefixes unreachable, VPN gateways "connecting" forever — look exactly like a local problem. The reflex is to bounce the circuit, check the router, and then open a ticket with the ISP. I've watched teams burn an hour on their own gear while the provider's status page already knew the answer. Every minute spent rebooting your router during a provider outage is a minute you weren't updating stakeholders with the real story.
Trap two: you trust the data plane and skip the management plane. This incident hit both, but the management-plane failures lingered after connectivity recovered. If your runbook ends at "tunnels are up, close the bridge," you miss the failed failovers, the half-applied firewall rules, and the VPN gateways running without redundancy. Recovery isn't over when the green light comes back. It's over when redundancy is verified back to baseline.
The checklist: is it us or the cloud?
This is the reusable part. Run it in order, write down what you find at each step, and stop troubleshooting your own gear the moment step one confirms the provider.
- Check the provider's status page and your Service Health dashboard before you touch anything. Azure status for the public view, Service Health in the portal for your tenants. This is step one because it's the cheapest and it rules out the biggest cause.
- Check breadth, not depth. Are other tenants in your org affected? Are services you don't route through your own gear affected (Teams, M365)? A provider event is wide; your broken router is narrow.
- Check your hybrid choke points in the portal, not from on-prem. ExpressRoute circuit status, VPN gateway connection status, BGP peer status — look at what Azure says about its side before you drive anywhere.
- Now check your side — the boring way. On-prem router CPU and logs, ISP circuit status, your BGP announcements (are you still advertising what you think you're advertising?), and the firewall logs on your end for the tunnel.
- Post an internal note early, with what you know. "We are seeing Azure connectivity degradation affecting ExpressRoute; provider has acknowledged on status page; investigating scope" beats silence every time. Timestamp everything.
- After recovery: verify redundancy, not just connectivity. This incident specifically degraded VPN gateway redundancy. Confirm gateways are back to active/active or your baseline, ExpressRoute circuits show both primary and secondary paths, and no management operations failed silently while the plane was down.
- Do the post-incident review while it's fresh — even for a provider outage. You didn't cause it, but your response time, your stakeholder updates, and your redundancy gaps are yours. File the PIR. Note what your runbook got right.
The boring lesson worth keeping
Every ops team has a change window. Microsoft's change window broke 18 regions of hybrid connectivity. You have exactly zero control over that, and no amount of SLA credits changes your users' evening.
What you do control is redundancy design and the first fifteen minutes. ExpressRoute with a single provider and a single circuit is a single point of failure wearing a premium SKU. Site-to-site VPN as a documented fallback, circuits from two providers into different peering locations, and — the cheapest one — a printed runbook that starts with the provider's status page instead of your router's console.
Cloud outages feel like acts of god. They're usually acts of maintenance. Plan like the provider's patching can take you down, because last night it did.
Sources
- Azure maintenance mess disrupts hybrid clouds, VPNs, cloudy VMware services — The Register, October 1, 2026
- Azure maintenance disrupts VPN gateways, ExpressRoute, and VMware services across 18 regions — lavx.hu, October 1, 2026
- Microsoft Azure Services Experienced Global Disruption — Wisevoter, October 1, 2026



