The Thinking Layer

Interactive explainers on AI

How AI actually works, one number at a time.

Models, infrastructure, research and real-world AI, explained with real numbers and visuals you can play with.

Signal & Weight · Part 1

One token through a 2.4-trillion-parameter model

Every word an LLM writes costs a full trip through its weights. I followed one prompt through Qwen3.8-Max to see what that trip actually looks like, in bytes, GPUs and memory.

Compare the models

Interactive · 10 open-weight models

Inside Open Models

Qwen3.8-Max, Kimi K3, DeepSeek V4, GLM-5.3, gpt-oss, Gemma 4 and more. Pick one and watch the same prompt travel through its real layers, experts, GPU memory and KV cache, computed from each model's published config.

Smallest
1 GPU
gpt-oss, Gemma 4, Mistral Small 4
Largest
16 GPUs
Qwen3.8-Max, Kimi K3
Cache spread
12×
DeepSeek V4 Pro vs Qwen at 128K tokens
DeepSeek V4 Pro's 384 experts, with six lit up by the router

Explore by topic

Series

8 parts · signalandweight.com

Signal & Weight

Where model weights live, how they move, and what they cost: one real model followed from storage to GPU memory to the next token.

Tools

2 interactives

All interactives

Inside Qwen3.8-Max, the nine-chapter journey of one token, and Inside Open Models, the same journey for ten models side by side.