One token through a 2.4-trillion-parameter model
Every word an LLM writes costs a full trip through its weights. I followed one prompt through Qwen3.8-Max to see what that trip actually looks like, in bytes, GPUs and memory.
Most conversations about AI infrastructure start with GPUs. How many, which generation, how much they cost. Almost nobody starts with the thing the GPUs spend most of their time waiting on: the weights.
So I wanted to see it for myself. I took one prompt, “Why is the sky blue?”, and followed it through one real model, Qwen3.8-Max, from the moment it’s typed to the moment the next word comes out. No hand-waving. Just the model’s published config file and arithmetic you can check.
Two numbers frame everything that follows. Qwen3.8-Max has about 2.42 trillion parameters. For any single token, it uses about 95 billion of them. Hold onto both.
Your words never reach the model
The model doesn’t see “Why is the sky blue?” It sees 14 integers.
A tokenizer cuts the prompt into pieces from a vocabulary of 248,320 entries, and Qwen’s chat template wraps it in special tokens that mark who is speaking. Everything downstream works on those numbers, not on letters.
The first weights are a lookup table
The first weights the model touches are the embedding matrix: 248,320 rows of 8,192 numbers each, about 2.03 billion parameters. And there’s no math here. Token 7 simply selects row 7.
Stack those rows and the prompt becomes a 14 × 8,192 grid called the residual stream. That grid is what travels through the rest of the model.
92 layers, three-to-one
Qwen3.8-Max stacks 92 layers in a rhythm: three layers of Gated DeltaNet, then one layer of full attention, repeated 23 times. Each layer adds its result to the residual stream instead of overwriting it. The vector gets refined on the way up, not replaced.
Two very different ways to remember
This is the design choice I find most interesting in this model.
The 23 full-attention layers keep a key and a value for every token they’ve seen, and compare each new token against all of them. To keep that affordable, 64 query heads share just 4 key/value heads, each 256 numbers wide.
The 69 DeltaNet layers skip the lookup entirely. Each one folds the new token into a fixed 128 × 128 memory matrix per head, with a gate that lets old information fade. That memory is the same size whether the conversation is ten tokens long or two hundred thousand.
One remembers everything, exactly. The other remembers a compressed summary, cheaply. Qwen mixes them three to one.
512 experts, and only 10 wake up
Every layer ends in a Mixture-of-Experts block: 512 small expert networks plus one shared expert. A router scores all 512 against the token and keeps the top 10. Each expert is about 50.3 million parameters (8,192 → 2,048 → 8,192), and only the chosen ones run.
That’s where the gap between 2.42 trillion and 95 billion comes from:
| Part of the model | Parameters |
|---|---|
| Routed experts (512 × 92 layers) | 2,370.8 B |
| Gated DeltaNet layers | 30.2 B |
| Full-attention layers | 9.6 B |
| Shared experts | 4.6 B |
| Embedding and output head | 4.1 B |
| Total | ≈ 2,419.8 B |
Almost 98% of the model sits in experts that any given token skips.
It doesn’t fit on one server
This is where it gets physical.
In FP8 (one byte per parameter), the weights alone take about 2.42 TB. An 8-GPU B200 server has 1.44 TB of GPU memory. Even an 8-GPU B300 server, at 2.3 TB, falls short. So one copy of this model has to span two servers: 16 GPUs holding about 151 GB of weights each, with roughly 29 GB per GPU left over for everything else.
Everything else, it turns out, is a lot.
The cache that grows with every token
Only the full-attention layers keep a cache that grows as the conversation gets longer. Its size per token comes straight from the config:
The DeltaNet layers hold a fixed 0.58 GB of state per conversation instead. The two break even at
and past that, the growing cache dominates. At the full 262,144-token context it reaches 24.7 GB per conversation. If all 92 layers used full attention, it would be 98.8 GB, four times as much.
That’s the real payoff of the three-to-one design. It isn’t about model quality on a benchmark. It’s about how many conversations fit on a GPU.
Then it does it all again
The last layer turns the vector into 248,320 scores, one per vocabulary entry. Softmax turns them into probabilities, one token gets picked, and it’s appended to the sequence. Then the whole trip starts over for the next word.
Here’s the part that surprised me most when I ran the numbers. Each of those trips reads about 95 GB of active weights from GPU memory to do roughly 190 billion floating-point operations. For a single conversation, the GPU spends far more time fetching weights than doing math.
That’s why inference servers batch many conversations into every step. And with 256 tokens in a step, the chance that a given expert is needed by at least one of them is
So in practice, nearly the whole 2.4 TB gets read every step. The trick is that the cost is shared across 256 tokens instead of one.
Same trip, ten different designs
Qwen3.8-Max is one set of choices. After finishing it I ran the same arithmetic on nine other leading open-weight models, from their own published configs, and the thing that jumped out wasn’t parameter count. It was memory.
Nine of the ten are Mixture-of-Experts models, and most touch only 3 to 6% of their weights per token (Nemotron is closer to 10%). Gemma 4 is the exception: dense, so every weight works on every token. Where they really split is in two decisions: how many bytes each weight takes, and how much memory each token of conversation costs.
| Model | Total · active | Weights | B200s | Cache at 128K |
|---|---|---|---|---|
| Kimi K3 | 2.8T · 104B | 1.60 TB FP4* | 16 | 4.1 GB |
| Qwen3.8-Max | 2.4T · 95B | 2.42 TB FP8 | 16 | 12.9 GB |
| DeepSeek V4 Pro | 1.6T · 49B | 0.87 TB FP4* | 8 | 1.1 GB |
| MiMo-V2.6-Pro | 1.02T · 42B | 0.55 TB FP4* | 4 | 6.8 GB |
| GLM-5.3 | 753B · 41B | 0.75 TB FP8 | 8 | 11.8 GB |
| DeepSeek V4 Flash | 284B · 13B | 154 GB FP4* | 2 | 0.7 GB |
| Nemotron 3 Super | 120B · 12B | 240 GB BF16 | 2 | 1.2 GB |
| Mistral Small 4 | 119B · 6.5B | 119 GB FP8 | 1 | 3.0 GB |
| gpt-oss-120b | 117B · 5.1B | 65 GB FP4* | 1 | 4.8 GB |
| Gemma 4 31B | 31B dense | 61 GB BF16 | 1 | 6.2 GB |
Two things in that table surprised me.
Kimi K3 is bigger than Qwen3.8-Max but lighter. It has 16% more parameters and still needs about a third less GPU memory, because its experts ship in 4-bit instead of 8-bit. That’s quantization deciding the hardware bill before anyone runs a single benchmark. It gets its own post later in this series.
DeepSeek V4 Pro’s cache is about 12 times smaller than Qwen’s. At 128K tokens, a Qwen conversation holds 12.9 GB of cache; DeepSeek V4 Pro holds about 1.1 GB, while being a 1.6-trillion-parameter model itself. DeepSeek compresses tokens 4× or 128× before caching them. Qwen swaps three of every four attention layers for DeltaNet. Kimi uses linear attention plus MLA. Nemotron leans on Mamba-2 and keeps only 8 attention layers out of 88. Different routes, same goal: fit more conversations on each GPU.
You can watch each of these designs yourself. Pick a model from the menu and the same nine chapters redraw with that model’s real layers, experts, GPU layout and cache curve.
FP4* means the experts are 4-bit while the rest stays at 8 or 16 bits. B200 counts assume about 153 GB usable per GPU for weights, with room left over for cache. Cache sizes are per conversation at 16-bit and include each model’s fixed state.
Where this series goes next
This is Part 1 of Signal & Weight. From here I want to follow these weights further: what they actually store, how quantization changes how many GPUs you need, how they get from storage onto GPUs in the first place, and what the KV cache and training checkpoints really cost. Every part uses this same model, so every number connects back to this one.
Sources: the published config.json for Qwen3.8-2.4T-A95B and the Qwen3.8 release, plus each comparison model’s published config on Hugging Face (linked inside the explorer). Parameter and memory figures are computed from the configs; GPU counts are illustrative.
Next in Signal & Weight: What weights actually store: compression, not a database (coming soon)