Hardware failure in large infra is not a question of “if”, but a question of “when” ⚡
While building a service that spans 100s of servers, never build with an assumption that the servers will remain up 24x7; instead, be pessimistic and assume failures at and after every single line of code.
Whatever that can go wrong, will go wrong. This is the secret of designing great systems.