
The kitchen

When running language models, expensive AI chips are often boring. To understand why, look inside a restaurant kitchen using this analogy:
- Ingredients represent the model weights and AI knowledge stored firmly in the warehouse.
- Guest orders represent tokens. They arrive from the outside and are processed step-by-step, recorded in heavy “stone notepads”.
Who Does What in the AI Kitchen?

Here is how a typical model operates on modern “Accelerator Processing Units” (APUs):
| Kitchen Role | Hardware Term | Plain English Explanation |
| 140 kg ingredient sacks | Model weights (VRAM) | The ingredients sit immovably in 140 kg sacks inside the warehouse. |
| Orders | Prompt / Tokens | The guests’ order needs to be applied to the ingredients. |
| Notepad clipboard | KV Cache (Context) | Previous orders. Every new order adds to the thickness of the notepad. |
| Waiter with a cart | Memory bus (Bandwidth) | Wheels ingredients and orders to the chef. Speed limit: 3.35 Terabytes (= grams) per second. |
| Chef cuisine with a knife | Compute cores (Compute) | Processes ingredients based on the orders: 989 trillion operations per second. |
The Kitchen Law

The ratio between the chef’s speed and the waiter’s carrying capacity determines utilization:
Ideal Ratio (I) = Trillion operations of the chef / Trillion grams the waiter can transport = 989 / 3.35 = 295 operations per gram of ingredient
- The Rule: For every gram of delivered ingredient, the chef must perform at least 295 operations to reach full capacity.
- The Problem: If a recipe requires fewer operations, the chef stands with a knife drawn, waiting on the sweating waiter.
The Two Phases: Understanding vs. Preparing

In an unoptimized kitchen workflow:
- Phase 1: Understanding (Prefill): The guest places 20000 requests. The waiter fetches the 140 kg sack just once. The chef processes all 20000 requests at the same time. The knife blazes, and the kitchen runs at 100% capacity.
- Phase 2: Preparing (Decoding): The dish is prepared one request at a time. For each single request, the waiter must wheel the entire 140 kg sack out of the warehouse again. The chef performs just one cut per ingredient and waits again.
The Kitchen Menu-Price Paradox

- The Fallacy: “Short and few orders ease the kitchen’s workload”.
- The Reality: For every request, the waiter still has to wheel in 140 kg of ingredients. The chef makes only one cut per gram of ingredient.
- The Result: The chef operates at less than 1% utilization. While short orders save notepad paper, they make remarkably poor use of the expensive chef. The menu price doesn’t really change; there will always be a baseline cost per ordered item (token).
Four Fixes for the Waiter Bottleneck

Because orders are generated sequentially, the waiter’s running speed dictates output speed:
- Powdered instead of raw ingredients (INT4 Quantization): The model is compressed. The sack shrinks from 140 kg to 35 kg. The cart is four times lighter, and words appear four times faster.
- Batching large orders (Batching): The waiter brings the sack once, and the chef serves 32 guests simultaneously. Utilization increases 32-fold.
- Shared notes (GQA – Grouped-Query Attention): Multiple chefs share a single clipboard. This saves up to 87.5% of note weight.
- The nimble kitchen apprentice (Speculative Decoding): A small sous-chef pre-chops 5 requests on a hunch. The main waiter brings the 140 kg sack once, and the head chef approves all 5 words in a single sweep.
Modern Kitchen Systems

| Technology | Kitchen Analogy | What Actually Happens |
| Mamba / SSM | The perpetual stock/reduction | Instead of piling up notes, the chef reduces the conversation into a lightweight 5 kg stock. Memory no longer grows linearly. |
| RAG & Vector DB | The archive pantry | Giant recipe books stay in the cellar. The waiter uses a flavor profile to fetch only the 3 most relevant summaries. |
| Mixture of Experts | Specialized stations | 16 specialized sacks in the warehouse (500 kg total). For a dessert order, the waiter fetches only 40 kg of ingredients. |
| Multimodality | Cooking with all senses | Image and audio orders are translated into the same format as text orders. The chef always sees orders in the same layout. |
| Reasoning / CoT | Scratchpad prep | The waiter pre-structures and breaks down complex orders for the chef. |
| Tool Use & ReAct | Specialized appliances | The chef uses dedicated tools, such as Python scripts, to prepare dishes more efficiently. |
Modern AI does not get faster simply by making the chef chop more frantically. Peak performance comes from lighter ingredients, compact reductions, smart specialized tools, and deliberate planning before serving.