Skip to visualization
T
Transformer
Atlas
01
Introduction
02
Tokenization
03
Embeddings
04
Positional encoding
05
Query · Key · Value
06
Self attention
07
Softmax
08
Attention matrix
09
Multi-head attention
10
Feed-forward network
11
Residual connections
12
Layer normalization
13
Encoder stack
14
Decoder stack
15
Masked self attention
16
Cross attention
17
Vocabulary projection
18
Autoregressive generation
19
Complete pipeline
20
Interactive playground
TRANSFORMER
ATLAS
LIVE SIMULATION
Frame 020
Embedding lookup
cat
1
2
3
4
5
6
7
8
+ positional signal
sin/cos(2)
ACTIVE REPRESENTATION
“cat”
token id 1137 · position 2
Strongest context
mat
36.3% attention · position encoded
Click a token to move camera focus
0.5×
1×
1.5×
2×
Chapter 01
14% through scene
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
LAB CONTROLS
Make the model yours.
Input sentence
The cat sat on the warm mat
7 tokens · updated live
Positional encoding
Preserve token order
Attention heads
4
Encoder layers
6