Certificate Lifecycle Management

Certificate Expiry Outages: Real-World Case Studies and What CLM Could Have Prevented

Executive summary — Certificate expiry is the most avoidable outage in enterprise IT, and also one of the most persistent. The certificate does exactly what it was designed to do — it stops being trusted on a date that was known months in advance — and yet the world's most sophisticated engineering organisations still get caught. Looking at what actually happened in a few well-documented incidents is more instructive than any warning, because each one maps cleanly to a specific control that would have caught it. This is a case-led argument for certificate lifecycle management as an uptime discipline rather than a compliance chore.

The pattern in every certificate outage is identical. A certificate that something depends on expires. The dependency was not on anyone's radar, either because nobody knew the certificate existed or because the renewal reminder went to a person who had changed roles. Service fails, often in a way that looks like a much scarier problem, and engineers lose hours ruling out attacks and infrastructure faults before someone thinks to check a validity date. The fix takes minutes; the diagnosis takes the outage.

The Incidents Everyone Remembers

In late 2017, LinkedIn suffered outages on country subdomains after an SSL certificate lapsed, leaving parts of the site unreachable. In February 2020, a global Microsoft Teams outage was traced to an authentication certificate that Microsoft had failed to renew, taking down collaboration for organisations worldwide for hours. In December 2018, an expired certificate inside Ericsson software knocked out mobile data for tens of millions of subscribers, including O2 customers in the UK and SoftBank users in Japan, in one of the largest certificate-driven outages on record. Around 2020, Spotify saw podcast delivery disrupted by an expired certificate. Different companies, different stacks, the same root cause.

Mapping Each Failure to a Control That Would Have Caught It

The LinkedIn subdomain case is a discovery failure: the certificate served real traffic but was outside whatever inventory the team was watching. Continuous discovery across all domains and subdomains would have surfaced it. The Teams case is a renewal-ownership failure: a critical certificate whose renewal depended on human memory rather than an automated pipeline. The Ericsson case is a dependency-visibility failure: a certificate buried inside software, invisible to the operators who ultimately felt its expiry. And the Spotify case is a monitoring failure: no alert fired with enough lead time to act. Each maps to a distinct lifecycle capability, and no single control would have caught all four — which is exactly why lifecycle management has to be complete rather than partial.

Reported Incident Root Cause Pattern Preventing CLM Control
LinkedIn subdomains, 2017 Certificate serving live traffic but outside inventory Continuous discovery across all domains and subdomains
Microsoft Teams, Feb 2020 Renewal dependent on human memory Automated renewal with owner-independent pipelines
Ericsson / O2 / SoftBank, Dec 2018 Certificate hidden inside software, invisible to operators Dependency mapping and full-stack certificate visibility
Spotify podcasts, c. 2020 No alert with usable lead time Expiry monitoring with escalating, early warnings

Incidents summarised from widely reported public accounts; details serve to illustrate certificate lifecycle failure modes rather than to attribute blame to any organisation.

Want to know which of your certificates could take a service down next quarter? CertiNext CLM discovers, monitors and renews across the whole estate.

Why the Risk Is Rising, Not Falling

Two forces are compounding this problem. First, the number of certificates in a typical enterprise is growing quickly as workloads, containers and services multiply, and the more certificates there are, the more likely one is forgotten. eMudhra's analysis of certificate sprawl sets out why this sprawl, not any external attacker, is the leading cause of self-inflicted downtime. Second, certificate lifetimes are shrinking. The CA/Browser Forum has agreed a phased reduction in maximum TLS certificate lifetimes to 200 days from March 2026, 100 days from March 2027 and 47 days from March 2029. Shorter lifetimes mean many more renewal events, and every manual renewal is another chance to miss one.

What 'Good' Looks Like Operationally

An organisation that has taken certificate expiry off the table can answer three questions instantly: how many certificates do we have and where, when does each expire, and what breaks if it does. It renews automatically wherever the endpoint supports it, so a human missing an email is not a single point of failure. It monitors the exceptions that cannot yet be automated with alerts that escalate well before the deadline. And it treats every new service as issuing a certificate that must join the inventory from day one, rather than as a surprise discovered during the next outage. The practical mechanics are covered in eMudhra's guide to automated certificate renewal with ACME.

The bridge to identity is worth naming, because certificates underpin far more than websites. They authenticate services, sign code and secure the machine-to-machine traffic that a modern estate runs on, which is why certificate hygiene sits alongside identity and access management as part of the same trust fabric. An outage from an expired certificate and a breach from a mismanaged credential are, at root, the same governance gap seen from two angles.

TAKE CERTIFICATE EXPIRY OFF YOUR OUTAGE LIST

eMudhra will inventory your certificate estate, flag the expiries that threaten uptime, and automate renewal before the next lifetime cut lands. Explore CertiNext CLM or speak to a CLM architect.

CertiNext Editorial
About the Author

CertiNext Editorial

CertiNext Editorial represents the collective voice of CertiNext, delivering expert insights on PKI modernization, crypto-agility, and the future of machine identity. Our team of PKI architects, security engineers, and digital trust specialists curates practical, in-depth content to help enterprises manage certificates at scale, eliminate outages, and prepare for the post-quantum era with confidence

Ready to Try?

Talk to our team about how eMudhra can help secure your digital workflows with PKI, eSignatures and identity solutions.

Connect with sales