Sissi Wang

neural networks cram lots of info into small numbers of neurons, so each neuron contain…

https://t.co/wrZtbQUzd2 neural networks cram lots of info into small numbers of neurons, so each neuron contains overlapping meanings a sparse autoencoder learns a BIGGER dictionary where each entry means one clean thing, but only a handful of entries activate at a time. this paper does smth rly cool: instead of penalising activations to make them sparse (which distorts them), just keep the top k and zero the rest. also mentioned metrics on if your dictionary's actually good! good stuff to read