Picking the right primary key is one of the most important design decisions of your schema; because an optimal key helps you avoid multiple unnecessary secondary indexes.
Here’s the schema of one of the production tables. The table
- holds 82GiB data
- has a primary key
- has no secondary indexes
- this table alone serves 2500 qps
- with a p90 response time of 3 ms
Apart from answering range queries efficiently, I am also powering slicing/dicing on one of the attributes. All of these with no secondary indexes, just a good old primary key.
Can you deduce the primary key of the table, and justify it with a reason in the comment? Let’s see how this goes.
Your design decision and reason will help others form intuition and understand the internals of the database. So do drop a note.
⚡ I keep writing and sharing these engineering nuggets, so if you are keen on learning them, follow along.