The Anatomy of a Production Outage
Walking through a real outage from the first alert to the last line of the post-mortem — what happened, what slowed the response, and what changed afterwards.
Articles on engineering, systems design, startups, and lessons from the trenches.
Walking through a real outage from the first alert to the last line of the post-mortem — what happened, what slowed the response, and what changed afterwards.
A postmortem is not a punishment. It is the mechanism by which incidents convert into system improvements.
Being on-call is a craft. There are habits that make it sustainable, and habits that make it miserable.
The CI/CD landscape has never been richer — or more overcomplicated. GitHub Actions covers 90% of real-world needs.
An application without monitoring is one whose failures you will learn about from your users.
Boring deployments are the goal. Here is the Git branching strategy and deployment discipline that makes deployments a non-event.
The documentation tells you how to build images. It does not tell you about the production realities.
Manual deployments are not a rite of passage. They are a liability.