wrote an essay on how LLM inference works

Arpit Bhayani

Arpit Bhayani

Nov 22, 2025 • 1 min read


wrote an essay on how LLM inference works. I spent about a week trying to build a clear understanding, and this is my distilled version of it. It covers the full inference journey: embeddings, attention, KV caching, quantization, and more.

I am no expert, but if you are looking for a starting point, this might help. Give it a read.

Arpit Bhayani

Principal Engineer II at Razorpay - building Agent Studio, Ex-staff engg at GCP Memorystore & Dataproc, Creator of DiceDB, ex-Amazon Fast Data, ex-Director of Engg. SRE and Data Engineering at Unacademy. I spark engineering curiosity through my no-fluff engineering videos on YouTube and my courses