tiny language model GPT
params
arch
vocab
ctx
step0
loss
steps/s0

training

corpus & tokenization
corpus - training data
loss
perplexity
|grad|
tokens0
sample generated while training
tokenization
vocab size100
vocabulary
learned merges
prompt, encoded
token frequency in corpus

transformer

value scale −max  0  +max

weights & biases

tensors
selected tensor

architecture

hyper-parameters
rebuild to apply

generate

input → output
input prompt
temperature0.85
top-k10
max tokens140
output 0 tokens
next-token