Writing an LLM inference loop
A ground-up look at tokenization, logits, temperature, softmax, and next-token sampling.
I build and operate reliable systems at scale, with a particular interest in SRE, Linux, distributed systems, and open source.
I’m a Production Engineer at Meta, and I write to understand the systems underneath everyday software.
A ground-up look at tokenization, logits, temperature, softmax, and next-token sampling.
Why LLM inference separates prompt processing from token-by-token generation.
Notes on purpose, hope, and what remains when circumstances strip almost everything else away.