Linear Mappings Act as Associative Memory, With Implications for Neural Network Training
A technical analysis explores what happens when a linear mapping is treated as a linear associative memory. The study finds that training beyond capacity does not erase old examples but instead gradually adds noise to stored associations, with recent examples recalled more accurately while older ones persist statistically. Below capacity, removing a training example may leave the weight vector unchanged unless weight decay is applied, which can shift the model to a lower-norm solution without losing remaining associations. These findings carry consequences for neural network initialization, weight decay, and stochastic gradient descent dynamics. The observations are also relevant to CCSLM architectures, where local experts can be interpreted as factorized associative memories.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in