We would like to read a transformer's virtual weights, but they include large confusing connections
|weight|
smalllarge
We can identify some interference weights using effectiveness: an estimate of how much each weight affects the function the model computes.
effectiveness
smalllarge
For a subset of weights, we can compute helpfulness: a measurement of how each weight affects the loss when ablated individually.
helpfulness
+
harmfulhelpful