r/LocalLLaMA 🤖 Ai 👁 0

A trained fast-weight memory: a 3M-param transformer installs never-trained rules at inference, forward-only — where test-time training transfers nothing (single RTX 3090, fully reproducible)

I'm an independent researcher (single self-funded RTX 3090). I just released a preprint (Zenodo for now — arXiv pending endorsement) on training a fast-weight memory bank: a small bank of vectors that the model writes wi

A trained fast-weight memory: a 3M-param transformer installs never-trained rules at inference, forward-only — where test-time training transfers nothing (single RTX 3090, fully reproducible)
📄

This source provides headlines only. Use the button below to read the complete article on the original site.

📰 Read the original article on r/LocalLLaMA

Originally published by r/LocalLLaMA. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.