Can SAEs Capture Neural Geometry?
manifolds explain SAE failures
Recall that lots of concepts are represented in activation space as manifolds (found via just varying the inputs, getting the activations, and then PCA’ing them down and seeing that they vary continuously)

SAEs are trying to capture this in a linear subspace. This explains a lot of the failures behind SAEs: 
Theoretically, SAEs could try to capture the manifold in one of these three ways 
- In practice, they just do dilution, which means that messily, each point on the manifold corresponds to a lot of feature directions, and each feature direction covers a sub-region of the manifold
You can try to reverse engineer the manifold from groups of correlated SAE features
- they should be correlated (in an Ising model sense) because, for instance, if you consider a feature for each day of the week, only one of these features should be on at each time
But they believe that you can decompose activation space into a superposition of low-dimensional manifolds 
- reverse engineering from SAEs is probably not the right way to go about finding these manifolds