r/LocalLLaMA
· Communities
A trained fast-weight memory: a 3M-param transformer installs never-trained rules at inference, forward-only — where test-time training transfers nothing (single RTX 3090, fully reproducible)
I'm an independent researcher (single self-funded RTX 3090). I just released a preprint (Zenodo for now — arXiv pending endorsement) on training a fast-weight memory bank: a small bank of vectors that the model writes with its own forward pass and reads as weights (each slot is expanded by a hypernetwork into a low-ran