PROJECT 01 / 18ML FUNDAMENTALSPYTHON

Research implementation

TorchZero.

A transformer, down to the derivative.

NumPyprimitive numerical backend
Reverse modeautodiff implementation
KV cacheincremental generation
01 / IDEA02 / SYSTEM03 / PLAYGROUND04 / DECISIONS05 / SOURCE
01 / THE IDEA

A closer look.

A neural network stack built without an automatic differentiation framework. TorchZero implements its own tensors, reverse-mode autodiff, transformer, tokenizer, optimizer, generation loop, and execution debugger over NumPy array primitives.

A training loop looks small because a framework hides its hardest contracts: how gradients accumulate through branches, how tensor shapes broadcast, and which activations must survive until backward. Rebuilding those contracts makes the whole path from bytes to a learned next-token distribution inspectable.

01

Own the gradient graph

Each Tensor stores its parents and a backward closure. A reverse topological traversal accumulates contributions into shared parents, checks gradient shapes, and releases graph edges unless retain_graph is requested.

02

Build the transformer

The decoder combines token embeddings, rotary positions, RMSNorm, causal multi-head attention, residual paths, and a feed-forward network. Incremental generation keeps detached keys and values for each layer.

03

Train, resume, inspect

The trainer coordinates shifted-token batches, cross-entropy, gradient clipping, AdamW, and a learning-rate scheduler. Checkpoint helpers and forward/backward traces expose the state behind a run.

04

Start at raw bytes

A byte-level BPE tokenizer learns frequent adjacent token merges. Dataset packing and a contiguous train/validation split keep the tokenizer and training pipeline in the same repository.

02 / UNDER THE SURFACE

From bytes to gradients

Follow training clockwise; inspect the generation branch below.

DRAG TO PAN · SELECT A NODE · + / − TO ZOOM

Read the architecture as text
  1. Byte-pair tokenizer — BPETokenizer starts with byte IDs and repeatedly merges the most frequent adjacent pair. Tie breaking is deterministic, and the merge vocabulary can be saved beside model checkpoints.
  2. Sequence packing — Packed token sequences become the input and one-position-shifted targets consumed by Trainer. A contiguous split separates training and validation before packed batches are sampled.
  3. Token embeddings — Embedding stores a trainable lookup table. Integer indices are checked for dtype and range before indexing the parameter tensor, so the resulting values remain connected to the autodiff graph.
  4. RMS normalization — RMSNorm divides by the square root of the mean squared activation plus epsilon and multiplies by learned weights. TransformerBlock applies it separately before attention and the feed-forward path.
  5. Query / key / value — Independent bias-free Linear layers map the normalized activation into queries, keys, and values. Each is reshaped into batch, time, head, and head-dimension axes before attention.
  6. Rotary positions — The attention forward pass applies rotary position information to queries and keys before moving heads into a separate axis. Values retain their projected representation.
  7. Causal attention — Queries multiply transposed keys, scaled by the inverse square root of head width. A lower-triangular mask blocks future tokens; softmax probabilities weight values and an output projection merges the heads.
  8. Residual block — TransformerBlock adds attention output back to the original activation, then adds the normalized feed-forward output. These additions create branching gradient paths through both the learned transform and the shortcut.
  9. Next-token loss — Trainer calls model.forward with shifted targets, then backpropagates the returned loss. This scalar connects the vocabulary prediction task to every differentiable operation in the forward graph.
  10. Tensor graph — The Tensor implementation supplies arithmetic, broadcasting, matrix multiplication, indexing, and shape operations. Every differentiable operation records enough local information to propagate an upstream gradient to its parents.
  11. Reverse traversal — backward seeds the root gradient, traverses nodes in reverse topological order, invokes each local backward closure, and sums contributions by parent identity. Shape mismatches and reuse of a freed graph raise errors.
  12. Clip + AdamW — The trainer clears prior gradients, differentiates the loss, clips the total gradient norm, advances the optimizer, and then updates the scheduler. AdamW applies decoupled weight decay directly to parameter values.
  13. Detached KV cache — Attention appends new keys and values to the per-layer cache and detaches stored tensors. The causal mask offset equals cached length minus query length, allowing a new token to attend to its prefix.
  14. Checkpoint state — Checkpoint routines serialize training state so the command-line interface can resume a run. The repository includes resume-equivalence tests; this page does not report a new execution of those tests.
  15. Execution debugger — The debugger captures forward information and traverses the backward graph to report tensor and gradient behavior. Attention tracing records Q/K/V shapes and the actual attention probability tensor.
03 / INTERACTIVE STUDY

Open the attention matrix

Change sequence length, head count, and width to see the causal attention footprint and a single-layer KV cache. Batch size is one.

CHANGE THE INPUTS

Illustrative shape accounting, not a benchmark or a running model. A real TransformerConfig requires width to divide evenly across heads; report incompatible combinations instead of implying they execute.

ILLUSTRATIVE MODELLIVE

04 / ENGINEERING CHOICES

Why it works this way.

01

Keep the numerical boundary explicit

NumPy supplies low-level arrays and primitive kernels. The tensor contract, local derivatives, graph traversal, neural modules, and training orchestration live in TorchZero itself.

02

Release graphs by default

Backward clears parent links and closures after use to release saved activations. Retaining a graph is explicit; accidentally differentiating through an already-freed graph is an error.

03

Separate measurement from execution

PyTorch appears in the benchmark comparison harness, not in the runtime implementation. Finite-difference tests and reference implementations provide different forms of correctness evidence.

05 / OPEN THE SOURCE

Trace it back.

Implementation details, examples, and project documentation.

Scope & limitations

  • This is a CPU-oriented NumPy implementation; it is not presented as a replacement for production accelerator kernels.
  • The diagram and attention lab explain source mechanisms. No new training run, source test suite, or performance benchmark was executed for this page.

Architecture and descriptions reflect the linked repository snapshot. The playground explains a mechanism; it does not execute the repository or report measured performance.