Some internal tools at cloud companies like AWS, GCP, and Azure are some of the most highly available systems ever designed ⚡
Tools like internal ticketing systems, alerting systems, monitoring stacks, collaboration tools, bastions, etc are typically meant to be used by internal engineers and be productive, so what’s the big deal?
These tools need to be designed in such a way that even when the entire cloud infrastructure is facing an outage, these tools still need to function; because during an outage, AWS engineers cannot say that we cannot debug because AWS is down.
They still need to collaborate, communicate, monitor, debug, and send out alerts, no matter what. Hence, even in internal tools, there are various tiers and the top-tier services are built with a really high degree of availability and fault tolerance.
Using a cloud is easy, but building one is pretty hard.
⚡ I keep writing and sharing these engineering nuggets, so if you are keen on learning them, follow along.