The Great Un-monolithing
An event-driven microservices architecture built on Kubernetes. More buzzwords please.
The Post-Gazette's publishing stack had been accreting since roughly 2016. One enormous application that did everything and feared nothing, least of all my weekends. Deploys were [ how deploys worked, and how scary ], and every change carried the whole system's risk.
The constraints were plain. One engineer. Bare metal on premises, no cloud autoscaling to hide behind. And a paper that publishes every single day, so there was no maintenance window big enough to swap the stack in one move.
Services stay small, own their data, and talk through RabbitMQ instead of calling each other directly. Varnish absorbs read traffic at the edge, so a couple million visits a month never becomes a couple million origin hits. MetalLB hands out real IPs, because bare metal does not come with a load balancer.
The rule that made 29 services survivable for one keeper: every service looks the same. Same repo layout, same base image, same health checks, same way onto and off the bus. The novelty budget goes to the problem, not the plumbing.
The monolith came apart one capability at a time. Peel a function off, stand it up as a service on the bus, route traffic to it, watch it for a while, then delete the old code path. Repeat 29 times. The first peel was [ the first service extracted, and why it went first ].
Order mattered more than speed. Low-risk, high-annoyance pieces went first, to prove the pattern. The scary ones went last, once the platform had earned trust.
Events everywhere is worth it for the decoupling, and you pay for it in debugging: tracing a story across queues is harder than reading a stack trace. [ how you trace across the bus today ]
The mistake worth admitting: [ the thing that actually broke once, and what it taught you ]