tiny language model GPT
params—
arch—
vocab—
ctx—
step0
loss—
steps/s0

training

corpus & tokenization
corpus - training data
—
loss—
perplexity—
|grad|—
tokens0
sample generated while training
tokenization
vocab size100
vocabulary
learned merges
prompt, encoded
token frequency in corpus

transformer

—
—
value scale −max  0  +max

weights

—
tensors
selected tensor

architecture

hyper-parameters
rebuild to apply
—

generate

input → output
input prompt
temperature0.85
top-k10
max tokens140
output 0 tokens
next-token