The Lab · Latent space

Inference lab

A trained model is a file of numbers. Serving it means moving those numbers through a chip fast enough, for enough people at once, at a price that makes sense. Six exhibits take the bill apart: why the first token costs differently from the rest, the cache that grows with every token, batching a stream of requests, rounding weights to fewer bits, letting a small model guess ahead, and a budget that ties it all together.

Everything runs in this tab. The attention decoder, the network you quantise and the speculative decoder are real, tiny, and seeded, so every visitor sees the same numbers. The hardware and model tables are approximate published peak figures and Llama-style shapes (shape only, not a specific product); real kernels reach about 30 to 70 percent of peak, so treat every time here as a floor, not a benchmark. Prices are illustrative, not quotes.