AI product Open source
Jacobian Lens (jlens) is a research and interpretability tool from Anthropic for examining what an internal language-model activation is disposed to make the model say. It linearly transports a residual-stream vector from any layer and position into the final-layer basis using an average input–output Jacobian computed over a text corpus, then applies the model’s own unembedding to produce ranked vocabulary tokens. The reference implementation fits lenses on open-weight decoder transformers, applies pretrained or newly fitted lenses, and renders an interactive layer-by-position view with top-token ranks, pinned-token tracking charts, and rank heatmaps. It is distributed as an Apache-2.0-licensed Python package and repository; the code is described as not maintained, and model weights and text corpora are not bundled.
1 use taken from transcripts — each links to the moment in the video.
A research tool for examining what an internal language-model activation is preparing the model to say. It projects activations into the output vocabulary, ranks tokens, and renders an interactive view across layers and positions.
1 in the library.