Future-state query path
The attention module predicts a latent state with an MLP and uses it for query projection, while keys and values come from the current state. A flag restores the ordinary query path.
Architecture research prototype
Explore attention, learned topology, and expert routing together.
An experimental PyTorch network that combines future-state query prediction, a learned token graph, and top-k expert routing. Its configuration and ablation paths make each mechanism separately inspectable.
Attention, graph aggregation, and expert selection offer different ways to mix information. GRAFT-Net puts these mechanisms inside one residual block so their interactions can be studied with explicit ablations.
The attention module predicts a latent state with an MLP and uses it for query projection, while keys and values come from the current state. A flag restores the ordinary query path.
Pairwise MLP scores produce a soft adjacency for diagnostics and a hard top-k mask for message aggregation. Sigmoid gates control the graph contribution.
Per-token utility scores choose top-k experts. The output includes selected indices and load fractions, with a dense FFN bypass for ablation.
The repository contains separate mechanisms, task heads, auxiliary losses, baselines, and ablation configuration, supporting experiments without hiding the paths being compared.
Trace tensor shapes through future queries, a learned graph, gated fusion, and experts. The training-target limitation is shown explicitly.
Adjust expert count, selected experts, and token-graph degree to inspect two different top-k operations on illustrative scores.
Illustrative routing only. Top-k is clamped to available experts. This is neither a trained model nor a speed or quality benchmark.
Predictive attention, latent topology, and routed experts each have a bypass. This creates explicit experimental comparisons rather than assuming every mechanism helps.
Topology constructs dense pair features before choosing neighbors. Expert routing selects sparse contributions, but selected experts still process full input tensors in this implementation.
The routing loss accepts gradient-utility targets, but the inspected task modules supply the routing scores themselves. The architecture should be presented as an experiment, not a demonstrated gradient-supervised advantage.
Implementation details, examples, and project documentation.
Inspected normalization, attention residual, gated topology fusion, and expert residual.
Inspected score prediction, top-k selection, full-tensor expert evaluation, and load accounting.
Inspected auxiliary outputs and the placeholder routing_targets assignment.
Architecture and descriptions reflect the linked repository snapshot. The playground explains a mechanism; it does not execute the repository or report measured performance.