Postmortem
API and dashboard unavailable
Summary
From 06:20 to 07:30 UTC, the API and the dashboard returned errors for every request. A certificate used between our load balancer and the application servers expired, so the load balancer refused to connect to them.
Impact
- 70 minutes of full downtime for the API and the dashboard.
- Requests failed with a 502 error. No data was lost, and scheduled jobs caught up afterwards.
- Checkout and the public website were not affected.
Timeline (UTC)
- 06:20 The internal certificate expires and requests start failing.
- 06:24 Monitoring alerts the on-call engineer.
- 06:51 The expired certificate is identified as the cause.
- 07:22 A new certificate is issued and rolled out.
- 07:30 Traffic is back to normal.
Root cause
The certificate was issued by hand two years ago, outside our automated renewal. Nothing warned us it was about to expire.
What we are changing
- Every internal certificate is now renewed automatically.
- We check certificate expiry daily and alert 30 days ahead.
- Our runbook lists every certificate and who owns it.