AWS Route 53 uses a technique called shuffle sharding to keep one bad shard from affecting 1/N customers. Here’s how it works…
If you split your instances into fixed shards, say 4 shards of 2 instances each, a bad request only hits one shard. So 1 in 4 customers gets affected if one shard has a problem. Better than nothing, but still a big blast radius.
Shuffle sharding does something simpler and smarter. Instead of fixed shards, each customer gets 2 random instances out of the pool. Two customers can share one instance, but it is very unlikely they share both. That randomness turns 4 possible shards into 56 combinations.
When something fails, the client retries the other instance in its shard. So if instance 5 is having a bad time, a customer whose shard is (3, 5) is still fine, because the client just falls back to instance 3.
Since clients retry a few times anyway, the odds of any one customer landing on a shard that is fully affected drop into the thousands. One noisy customer or one bad shard, instead of taking down a quarter of your traffic, ends up affecting almost no one else.
By the way, this is also the same reason a Bloom filter uses multiple hash functions instead of just one. More independent random choices, less chance they all collide.