Own the gradient graph
Each Tensor stores its parents and a backward closure. A reverse topological traversal accumulates contributions into shared parents, checks gradient shapes, and releases graph edges unless retain_graph is requested.
A neural network stack built without an automatic differentiation framework. TorchZero implements its own tensors, reverse-mode autodiff, transformer, tokenizer, optimizer, generation loop, and execution debugger over NumPy array primitives.
A training loop looks small because a framework hides its hardest contracts: how gradients accumulate through branches, how tensor shapes broadcast, and which activations must survive until backward. Rebuilding those contracts makes the whole path from bytes to a learned next-token distribution inspectable.
Each Tensor stores its parents and a backward closure. A reverse topological traversal accumulates contributions into shared parents, checks gradient shapes, and releases graph edges unless retain_graph is requested.
The decoder combines token embeddings, rotary positions, RMSNorm, causal multi-head attention, residual paths, and a feed-forward network. Incremental generation keeps detached keys and values for each layer.
The trainer coordinates shifted-token batches, cross-entropy, gradient clipping, AdamW, and a learning-rate scheduler. Checkpoint helpers and forward/backward traces expose the state behind a run.
A byte-level BPE tokenizer learns frequent adjacent token merges. Dataset packing and a contiguous train/validation split keep the tokenizer and training pipeline in the same repository.
Follow training clockwise; inspect the generation branch below.
Change sequence length, head count, and width to see the causal attention footprint and a single-layer KV cache. Batch size is one.
Illustrative shape accounting, not a benchmark or a running model. A real TransformerConfig requires width to divide evenly across heads; report incompatible combinations instead of implying they execute.
NumPy supplies low-level arrays and primitive kernels. The tensor contract, local derivatives, graph traversal, neural modules, and training orchestration live in TorchZero itself.
Backward clears parent links and closures after use to release saved activations. Retaining a graph is explicit; accidentally differentiating through an already-freed graph is an error.
PyTorch appears in the benchmark comparison harness, not in the runtime implementation. Finite-difference tests and reference implementations provide different forms of correctness evidence.
Implementation details, examples, and project documentation.
Q/K/V projections, rotary positions, causal masking, residual blocks, and detached cache handling.
Topological traversal, accumulation, shape checks, and graph cleanup.
Loss, backward, clipping, optimizer, scheduler, timing, and checkpoint flow.
Documents gradient checks, reference parity, cache parity, and the external numerical boundary.
Architecture and descriptions reflect the linked repository snapshot. The playground explains a mechanism; it does not execute the repository or report measured performance.