Postmortem
Sign-in failing for some customers
Summary
From 16:05 to 17:10 UTC, about one in four sign-ins to the dashboard failed. One node of our session store ran out of memory and stopped accepting new sessions. Customers who were already signed in were not affected.
Impact
- 65 minutes in which about 25% of sign-ins failed.
- Around 1,300 sign-in attempts failed. Most succeeded when retried.
- The API and existing sessions kept working.
Timeline (UTC)
- 16:05 Sign-in errors start rising.
- 16:11 The on-call engineer starts investigating.
- 16:32 One session store node is found out of memory.
- 16:49 The node is replaced and sign-ins recover.
- 17:10 Error rates have been normal for 20 minutes.
Root cause
After a configuration change a week earlier, expired sessions were no longer removed from that node. Its memory filled up slowly until it could not store new sessions.
What we are changing
- The expiry setting is now checked automatically on every node.
- We alert when any session store node passes 75% memory.
- The sign-in error rate now pages the on-call engineer directly.