Essay : Everything you need to know about handling node

Arpit Bhayani

Arpit Bhayani

Sep 07, 2021 • 1 min read


Essay #63: Everything you need to know about handling node outages in Master Replica setup.

This essay talks about the worse - nodes going down - impact, recovery, and real-world practices.

The outline of the essay is:

  • Nodes go down, and it is okay

  • Handling Replica Outages

  • Handling Master Outages

  • Manual way

  • Automated way (through Leader Election)

  • How Master outages are addressed in real world

  • Understanding the Passive Master

The long-form of this snippet where I discuss the outages in detail can be found at: https://lnkd.in/dS6wwDWv.


To date, I have written 63 articles on Distributed Systems, System Design, Advanced Algorithms, and Python Internals.

Right now, I am running a series on Distributed Systems and System Design. ✨

1900+ people have subscribed to my newsletter.


If you seek to become a better engineer and learn System Design the right way, you can join the waitlist for my November batch.

You can find the week-by-week curriculum and topics, benefits, testimonials, and other details: https://lnkd.in/dtBk7eE.

Arpit Bhayani

Principal Engineer II at Razorpay - building Agent Studio, Ex-staff engg at GCP Memorystore & Dataproc, Creator of DiceDB, ex-Amazon Fast Data, ex-Director of Engg. SRE and Data Engineering at Unacademy. I spark engineering curiosity through my no-fluff engineering videos on YouTube and my courses