Testing the global workspace at open frontier scale: the Jacobian lens on DeepSeek-V4-Flash and GLM-5.2

A community replication of the Jacobian lens on DeepSeek-V4-Flash and GLM-5.2: structural signatures, causal interventions, a preregistered selectivity sweep, and a comparison of two DeepSeek post-training releases.

TL;DR — We bring Anthropic’s Jacobian lens to DeepSeek-V4-Flash and GLM-5.2, two open-weight models with hundreds of billions of total parameters. To our knowledge, this is the first community replication of the method and its global-workspace experiments at this scale. The port supports DeepSeek’s four-stream residual architecture and fitting through GLM’s native FP8 kernels.

Code and experiment scripts · Fitted lenses: DeepSeek-V4-Flash-0731 · DeepSeek-V4-Flash preview · GLM-5.2

1. What this replication tests

Verbalizable Representations Form a Global Workspace in Language Models introduces the Jacobian lens as a way to read internal activations in vocabulary space. The paper argues that verbalizable representations occupy a middle region of the network with several properties associated with a global workspace: limited capacity, an ignition-like transition, causal importance for deliberate reasoning, and relative separation from automatic processing. Its experiments use Claude models, whose original checkpoints are unavailable for independent replication.

Community work has made the method accessible on open models, including a Qwen port, a GPT-2/Qwen replication, an Apple-Silicon visualizer, and Neuronpedia’s fitted lenses. These efforts provide useful tests across architectures and model sizes. Our contribution extends that work to open models at frontier scale, where both model capability and the computational requirements are substantially different.

We study DeepSeek-V4-Flash-0731 and GLM-5.2. Both have hundreds of billions of total parameters; GLM has 753B total and approximately 40B active per token. We use frontier scale to describe this size and their position in their developers’ flagship model lines, without assuming that their capabilities match the Claude checkpoints used in the paper. The two models also present different implementation challenges: DeepSeek has four parallel residual streams, while GLM uses a conventional single-stream residual but requires backward support for FP8 kernels.

The study asks two questions. Which of the paper’s operational signatures can we recover on these open checkpoints? And what does the lens reveal about differences between two DeepSeek releases described as sharing a pretrained base but differing in post-training? The main replication concerns DeepSeek-V4-Flash-0731 and GLM-5.2. The extension adds DeepSeek’s April preview; a smaller exploratory comparison also uses its base checkpoint.

The contribution includes both the replication and the infrastructure needed to make it possible:

Three kinds of evidence must be kept separate. A lens can read out a concept; an intervention can show that selected directions are causally important; and an intervention can be selective, disrupting one class of behavior while preserving another. Our clearest result concerns the distinction between the latter two. J-lens ablations strongly affect multi-hop performance, but we do not find the reasoning/prediction separation reported for Claude within our tested implementation and parameter range.

2. From residual activations to vocabulary readout

The Jacobian lens transports an intermediate residual activation into the model’s final representation space, then applies the model’s final normalization and unembedding:

\[\operatorname{lens}_{\ell}(\mathbf h) = W_U\,\operatorname{Norm}\!\left(\widehat J_{\ell}\mathbf h\right).\]

The transport matrix is estimated from derivatives of downstream activations with respect to the source residual. Intuitively, it describes how changes at a layer propagate toward the output, averaged over a text corpus. Vocabulary tokens then provide an interpretable readout of the transported activation. The fitting procedure uses automatic differentiation and averaging, without optimizing a prediction loss.

Let \(G_{\ell}^{(i)}\) denote the Jacobian contribution of prompt \(i\), including its source and target position reductions. The original estimator takes an arithmetic mean. Most of our experiments instead use a mean of per-prompt, per-layer normalized contributions:

\[\widehat J_{\ell}^{\mathrm{plain}} = \frac{1}{n}\sum_{i=1}^{n}G_{\ell}^{(i)}, \qquad \widehat J_{\ell}^{\mathrm{norm}} = \frac{1}{n}\sum_{i=1}^{n} \frac{G_{\ell}^{(i)}}{\lVert G_{\ell}^{(i)}\rVert_F}.\]

These generally estimate different population quantities. Normalization reduces the influence of high-gain prompts, which we observed on DeepSeek. We therefore distinguish this choice from the architectural adaptations and test its effect on the main ablation result below. Appendix A gives the position reduction and implementation details.

Models and fitting

Checkpoint Layers Residual architecture Jacobian shape, target × source Fitted layers and corpus
DeepSeek-V4-Flash-0731 43 Four streams, width 4096 each 4096 × 16384 All layers at n=100; L19–39 also at n=1000
GLM-5.2 78 Single stream, width 6144 6144 × 6144 L34–71, n=100
DeepSeek-V4-Flash preview 43 Four streams, width 4096 each 4096 × 16384 L19–39, n=100

On DeepSeek, the source is the flattened four-stream residual and the target is the output of hc_head, which combines those streams. The relevant map is therefore rectangular. We generalized the reference library’s source and target dimensions while retaining the single-stream path. The logit-lens baseline also passes intermediate residuals through the final stream-combination module before decoding.

Choosing the collapsed target also keeps the number of backward batches proportional to 4096 output coordinates rather than 16384 flattened coordinates. The port includes tests for stream ordering, finite-difference agreement on a tiny V4 model, readout behavior, and compatibility with older square checkpoints. These are checks of the adaptation; they do not replace validation of the large-model experiments.

GLM’s geometry fits the square formulation, but its native FP8 matrix-multiplication kernels lack backward formulas. Dequantizing the full model would require approximately 1.5 TB, exceeding the available 1.17 TB of GPU memory. We registered activation-gradient formulas for three FP8 operations, keeping weights frozen and treating internal activation quantization as the identity during backward. This produces a straight-through Jacobian approximation. The n=100 band fit took 14.2 hours on the available eight-GPU system; the approximation has not been validated against finite-difference directional derivatives.

Lenses use WikiText sequences of up to 128 tokens, excluding the first 16 source positions and the final position from fitting. The task evaluations use the released synthetic prompt sets. Band statistics and next-token agreement use WikiText samples that overlap the fit corpus: the loader takes deterministic prefixes of the same stream. Those statistics are therefore descriptive measurements on the fitting distribution, not held-out evaluations.

What the metrics measure

Readout recovery asks whether a designated target appears as the lens’s top token at any scored layer. The evaluated layer range is specified for each comparison. For items with several intermediates, we average the recovered fraction within each item and then across items, following the set-specific readout positions and target variants. We retain the released name pass@1; it is a vocabulary recovery score, not a sampled-generation success rate. Token ranks are zero-indexed.

Multi-hop accuracy is first-token answer accuracy on 90 two-hop prompts. The selectivity follow-up strips trailing whitespace before tokenization. Next-token agreement measures how often an intervened model’s argmax matches the unablated model’s argmax, over 2,220 positions from 20 WikiText prompts. It is different from the band statistic called top-1 accuracy, which compares the lens prediction with the actual next corpus token.

Ablation removes the projection of each residual onto directions derived from that position’s top-\(k\) lens tokens, repeated at every selected layer. The subspace can change with both position and layer. Thus, “rank 10” describes each local edit, not a single rank-10 subspace shared across the network.

Anthropic released the lens implementation but not the causal intervention code. Our interventions implement the described procedures; numerical comparisons with Claude also span differences in evaluation protocols. Throughout this post, Claude values are published reference results rather than measurements from a shared implementation.

3. Structural evidence for a candidate workspace band

Layer statistics and geometry

On DeepSeek-V4-Flash-0731, the layer statistics suggest a band at L19–39. The lens’s next-token accuracy is 0.0059 at L19, rises to 0.1665 at L39, and reaches 0.2452 at L40. Effective dimensionality expands through much of this region before dropping sharply at L40. CKA provides complementary evidence: the layer geometry forms three blocks with boundaries near L19 and L40.

Four layer-wise statistics on DeepSeek-V4-Flash: next-token accuracy, excess kurtosis, excess autocorrelation, and effective dimensionality, with L19–39 highlighted.
Figure 1 (cf. paper Fig. 28): Layer statistics used to identify DeepSeek's candidate band. Effective dimensionality reaches 7.58 at L39 and falls to 1.29 at L40, while next-token accuracy rises toward the output. The WikiText sample overlaps the lens-fitting corpus; curves describe one corpus sample and have no uncertainty bands.

L19–39 corresponds to approximately 44–91% of depth, compared with the paper’s 38–92% on Sonnet 4.5. This is a similar relative location despite DeepSeek’s four-stream residual architecture. We use candidate workspace band for this empirically selected region; the label does not by itself establish all the functional properties of a global workspace.

For GLM, we fitted only the window projected from DeepSeek’s relative depth, L34–71. Its next-token accuracy increases from 0.0012 at L34 to 0.194 at L71. The window contains a compatible low-to-high transition, but this design does not independently locate GLM’s boundaries. Fitting layers outside the window is necessary to determine whether the onset is earlier or the offset later.

An ignition-like transition on DeepSeek

The ignition experiment interpolates between two token embeddings,

\[\mathbf e(\alpha) = \alpha\,\operatorname{emb}(A)+(1-\alpha)\,\operatorname{emb}(B),\]

and measures the relative readout strength of A and B. The statistic is reciprocal-rank share, \(\mathrm{RR}_A/(\mathrm{RR}_A+\mathrm{RR}_B)\), where \(\mathrm{RR}=1/(1+\mathrm{rank})\). It summarizes the two vocabulary ranks and should not be interpreted as a calibrated probability.

On DeepSeek, the 10–90% transition occupies 0.60 in \(\alpha\) at L19, becomes unresolved within a single 0.10 grid step at L24–33, and broadens to 0.70 at L39. At L30, the share changes from 0.00 to 0.05, 0.62, and 0.94 as \(\alpha\) increases from 0.3 to 0.6.

DeepSeek layer-by-interpolation heatmap showing a sharp change in reciprocal-rank share within the candidate band.
Figure 2 (cf. paper Fig. 29): The readout transition is sharpest inside the candidate band. Widths at the 0.10 grid resolution do not resolve the underlying transition shape. Early layers do not show the same threshold pattern; some track the mixture in the opposite direction.

This reproduces the qualitative ignition-like signature on DeepSeek. A finer interpolation sweep and additional token pairs would be needed to characterize its sharpness and consistency. The evidence presented here does not include an equivalent full-depth ignition measurement on GLM.

Arithmetic as a readout example

For calc: (4+17)*2+7=, DeepSeek’s intermediates first become the lens’s top token at L27 (21), L30 (42), and L36 (49). Their emergence follows the order of the calculation, within the candidate band. On a correctly answered GLM instance, the readout recovers 21 at L50 and 49 at L67; it does not reproduce all three DeepSeek milestones.

Ranks of the tokens 21, 42, and 49 across DeepSeek layers for one arithmetic expression.
Figure 3 (cf. paper Fig. 17): A single-prompt arithmetic trajectory. For logarithmic plotting, the figure uses rank + 1, so its top token is labeled 1 rather than the zero-based rank used in the text. The annotation “cleared” describes deterioration in vocabulary rank; it does not establish active removal of the underlying representation.

This example establishes that the lens can surface ordered intermediates in these cases. It does not measure how reliably arithmetic is represented or recovered. The broader six-set evaluation is mixed: on DeepSeek, the J-lens outperforms the logit lens on association and typo recovery, but trails it on multi-hop, multilingual, and order-of-operations recovery. The separate tuned-lens evaluation also has an uneven profile. Arithmetic therefore provides both a useful illustration and a limit on generalization from that illustration. Appendix C reports the complete comparison.

4. Causal effects and the selectivity test

The causal experiment asks whether J-lens directions matter for behavior. The stronger selectivity test asks whether removing them impairs multi-hop while leaving ordinary next-token predictions largely intact. We evaluate both outcomes under the same intervention.

The main ablation result

The following table uses the corrected multi-hop tokenization and a consistent agreement protocol for both models. The intervention is position-adaptive, \(k=10\), across L19–39 for DeepSeek and L34–71 for GLM.

Model Unablated multi-hop J-ablated multi-hop Next-token agreement under J-ablation
DeepSeek-V4-Flash-0731 74/90 (0.822) 26/90 (0.289) 0.523
GLM-5.2 76/90 (0.844) 29/90 (0.322) 0.498

Multi-hop accuracy falls by 48 items on DeepSeek and 47 on GLM. At the same time, roughly half of the unablated next-token argmaxes change. The two agreement values differ by 0.025; both are substantially below the paper’s reported agreement above 0.90 on Claude. Agreement measures prediction changes, however, not whether each changed prediction is incorrect.

Earlier versions of this post used the raw prompt strings and reported 53/90 → 9/90 on DeepSeek and 55/90 → 15/90 on GLM. Trailing whitespace depresses first-token scores on this set. The corrected protocol is used throughout this section; the earlier measurements and agreement-denominator correction are documented in Appendix B.

Random controls at comparable perturbation magnitude

A rank-matched random ablation provides an initial control, but equal rank need not imply equal perturbation magnitude. We therefore also subtract random unit directions scaled to a target fraction of the residual norm. That target is the J-ablation’s mean perturbation norm fraction measured on WikiText. We run five random seeds at each magnitude.

Model Intervention WikiText norm fraction KL(intervened ∥ unablated), nats Agreement Multi-hop accuracy
DeepSeek J-space, k=10 0.0351 2.548 0.523 0.289 (26/90)
DeepSeek Random, norm-calibrated 0.0351 0.105 0.896 0.827, mean over 5 seeds
GLM J-space, k=10 0.0990 3.404 0.498 0.322 (29/90)
GLM Random, norm-calibrated 0.0990 1.055 0.736 0.784, mean over 5 seeds

The random multi-hop counts are 76, 74, 74, 73, and 75 out of 90 on DeepSeek, and 72, 71, 73, 69, and 68 on GLM. Their means are 74.4 and 70.6 correct items. DeepSeek’s random mean remains near its 74/90 baseline; GLM’s is lower than its 76/90 baseline, but much closer to it than the J-ablated result.

The random-minus-J accuracy gaps are 0.538 and 0.462. They are 16.1 and 8.3 times the corresponding five-seed min–max ranges. These ratios describe variation across sampled random directions; they are not significance levels or uncertainty estimates over tasks.

The norm match also has a precise scope. It is a match to a WikiText mean, not to the local J projection at every layer and position. On the multi-hop prompts, the mean all-position perturbation fractions are 0.0387 versus 0.0351 on DeepSeek and 0.0978 versus 0.0990 on GLM for J and random respectively. The magnitudes are comparable, with a modest mismatch on DeepSeek. The scaled random edit is a norm-controlled perturbation rather than an orthogonal projection removal.

Multi-hop accuracy versus WikiText perturbation norm and next-token KL for J-space and three random-control constructions on both models.
Figure 4: Dose-response curves for the corrected multi-hop protocol. Horizontal-axis norm and KL measurements use WikiText; vertical-axis accuracy uses the 90 multi-hop items. Random-control bands show the minimum and maximum over five seeds. The arm labeled norm-matched uses the WikiText calibration described above. Sampling directions inside rowspace(J) constrains them to the lens's row space, without establishing that they lie on the activation manifold.

The wider curves show that sufficiently large random edits can also impair multi-hop. On GLM, the rank-varying random arm reaches mean accuracy 0.205 at a norm fraction of approximately 0.199. DeepSeek’s J-space curve is non-monotone in rank, with less damage at some larger ranks. Consequently, neither rank nor perturbation norm alone summarizes all intervention behavior. At the reported operating point, the controls nevertheless support a substantial dependence on which directions are perturbed.

Searching for reasoning/prediction separation

We preregistered the following criterion: an intervention is selective if next-token agreement is at least 0.85 and multi-hop accuracy is at most half of its same-run unablated baseline. The multi-hop cutoffs are therefore 37/90 on DeepSeek and 38/90 on GLM. The 0.85 threshold is more permissive than the paper’s reported agreement above 0.90.

The grid varies \(k\in\{1,2,5,10,20,50\}\), four layer ranges (full band and its early, middle, and late thirds), and fixed versus position-adaptive token selection. Adaptive selection is tested with and without excluding the unablated model’s top predicted token. Fixed selection uses a corpus-level token set without that per-position guard. This gives 72 grid treatments, plus one legacy lens-top-1 guard treatment, for 73 per model. The protocol and its amendments specify these choices before the grid runs; one amendment was recorded after observing the baseline tokenization discrepancy.

No treatment on either model meets the selectivity criterion. Selected DeepSeek configurations illustrate the tradeoff:

Configuration, k=10 Next-token agreement Multi-hop correct / 90
Full band, adaptive 0.523 26
Late third, adaptive 0.535 34
Middle third, adaptive 0.789 69
Early third, adaptive 0.876 74
Full band, fixed corpus-level selection 0.908 75

The early-band and fixed-subspace interventions preserve agreement and multi-hop together. The more damaging adaptive interventions reduce both. Across the complete grid, no configuration combines the required agreement with a reduction to half-baseline multi-hop.

Next-token agreement versus multi-hop accuracy for 73 treatment configurations per model, with no treatment inside either preregistered selectivity region.
Figure 5: The preregistered selectivity search. Each open-model treatment is compared with its own baseline-dependent cutoff. The Claude annotation summarizes the paper's reported behavior under its protocols; it is not a Claude measurement from this sweep. No tested treatment reaches either open-model selectivity region. Agreement of 1.0 means unchanged argmaxes on the scored positions, not necessarily unchanged output distributions.

Two further checks address specific implementation choices. First, excluding the model’s own top predicted token improves DeepSeek agreement from 0.523 to 0.594, with multi-hop at 40/90. This is a meaningful recovery, but it still does not meet the selectivity criterion. The guard reduces direct overlap with the predicted token; it does not protect every direction relevant to prediction.

Second, repeating the full-band, k=10 experiment with the plain-mean estimator gives:

Model Normalized-lens agreement / multi-hop Plain-mean-lens agreement / multi-hop
DeepSeek 0.523 / 26 of 90 0.559 / 31 of 90
GLM 0.498 / 29 of 90 0.477 / 31 of 90

The estimator changes the measurements, but all four conditions remain non-selective by our criterion. This check shows that removing normalization does not restore selectivity at this operating point. It does not establish equivalence of the estimators for all interventions or readout tasks.

Finally, we attempted a positive control on smaller open models. GPT-2 scored 8/90 unaided, below the preregistered feasibility gate of 0.30. Qwen3-1.7B, Qwen3-4B, and Qwen2.5-7B-Instruct passed that gate, but none yielded a selective configuration. Fixed rank also corresponds to different perturbation magnitudes across models: at k=10, the mean norm fraction was 0.035 on DeepSeek and 0.143 on Qwen3-1.7B. This complicates the comparison, especially where random controls also impair performance.

The positive-control search therefore does not demonstrate that our pipeline detects selectivity where it exists. Together, the experiments support direction-specific causal effects and a failure to find task selectivity in the tested configurations. They narrow several explanations for the difference from Claude, but do not eliminate all differences between implementations or establish a model-family cause.

5. Additional behavioral probes

Concept readout under suppression

The suppression experiment tests whether a named concept remains readable while the model answers an unrelated question. For each eligible concept, we compare instructions such as “Think carefully about elephant while answering: name a color” and “Do NOT think about elephant at all. Name a color.” We generate eight tokens and score a hit if the selected concept token becomes the lens’s top token at any generated position and band layer.

Reading generated positions avoids directly scoring the concept’s appearance in the input. We also check whether the generated text verbalizes the concept. Concepts must have an eligible single-token form, leaving 13 items on DeepSeek and 12 on GLM.

Model Focus Suppress Absent-concept condition Verbalized in focus
DeepSeek-V4-Flash-0731 12/13 0/13 Not measured 10/13
GLM-5.2 12/12 11/12 0/12 0/12

The suppression contrast is large on this small set: the descriptive Wilson 95% intervals are [0, 0.23] for DeepSeek and [0.65, 0.99] for GLM. Neither model verbalizes the concept in the scored suppression continuations. DeepSeek’s working focus condition establishes that the readout can detect the concepts under a different instruction, but it is not a substitute for the absent-concept control measured on GLM.

The result is a difference in detectable target-token readout during a short generated span. It does not establish that DeepSeek removes every representation of the concept, or that GLM cannot suppress it over a longer continuation. Different eligible items, chat formats, and searched layer counts also limit how precisely we can attribute the contrast to internal control mechanisms. Extending the comparison to a common concept set and matched negative controls would make it more informative.

Swap and broadcast

A swap edits the lens-derived direction associated with an intermediate concept and asks whether the immediate answer changes to the target answer. A broadcast test asks whether an edit affects an answer to a different downstream question that depends on that attribute. The latter requires a behavioral effect beyond making the injected concept readable.

Model or reference Swap: target answer among baseline-correct items Broadcast: target downstream answer
Claude, as reported in the paper 0.54–0.70 0.526
DeepSeek-V4-Flash-0731 23/53 (0.434) 40/192 (0.208)
GLM-5.2 25/55 (0.455) 30/80 (0.375)

These experiments retain the original raw-prompt scoring protocol, and use the \(\alpha=2\) operating point selected on the same items. Their swap denominators must not be replaced with the corrected ablation baselines: doing so would require rerunning the interventions and redefining which items were baseline-correct.

The swap estimates are close, but the sample sizes do not establish equivalence. Broadcast is higher on GLM in these runs; item-level Wilson intervals are approximately [0.28, 0.49] for GLM and [0.16, 0.27] for DeepSeek. Related items may share concepts or templates, so these intervals can understate uncertainty. The different denominators and the tuned operating point further limit a direct model-level comparison.

The two readouts nevertheless ask distinct causal questions: whether an edit changes the immediate answer, and whether another computation uses the edited attribute. Their observed pattern motivates studying those questions separately. GLM’s similar swap estimate also shows that DeepSeek’s four-stream architecture is not necessary for obtaining a swap rate below the paper’s reference range; it does not isolate the cause of either shortfall.

Experiential language and the story-quality control

Full-band ablation also changes the style and coherence of generated language. Each condition contains three experiential continuations and one story continuation. Each passage receives three binary rubric judgments. Experiential scores therefore average nine judgments across three passages; story-quality scores average three judgments about one passage. These are rubric scores, not success rates over nine independent generations.

Independent grader Condition Experiential rubric Story-quality rubric
Qwen2.5-7B Unablated 0.889 1.000
Qwen2.5-7B Full-band J-ablation 0.222 0.000
Qwen2.5-7B Random control 0.889 1.000
Qwen3.6-35B Unablated 0.667 1.000
Qwen3.6-35B Full-band J-ablation 0.222 0.000
Qwen3.6-35B Random control 1.000 0.667

The graders see passage text without intervention labels. Both identify deterioration under full-band J-ablation, including on the story-quality control. This is consistent with a broad effect on generation rather than an effect specific to experiential reports. With only three experiential passages and one story per condition, the result is exploratory. The earlier 0.778/0.333 baseline came from self-grading with the original story rubric; it should not be attributed to the independent Qwen2.5 grader.

6. Comparing two DeepSeek post-training releases

DeepSeek describes the April preview and July 0731 release as having the same architecture and size, with post-training redone. This provides a second comparison: how much of the lens’s measured behavior differs between two releases within one model family?

The checkpoint inventory supports part of the premise. Both releases have the same 67,612 main-stack tensor names and shapes, both ship MXFP4 experts, and all 48 weight-shard hashes differ. These checks establish matching architecture and different released weights; they cannot establish identical pretrained starting weights. We therefore treat this as a vendor-described same-base comparison.

Prompt formatting also requires care. The reasoning-effort names changed between releases: a prefix associated with preview’s max is associated with 0731’s high. We pinned the literal prefixes and verified identical encodings for the comparisons. Each checkpoint’s lens uses the same n=100 fitting recipe over L19–39.

Lens similarity and token overlap

A cross-checkpoint difference combines changes in the model with variability in the fitted lens. As a reference, we compare two lenses of 0731 itself, fitted on disjoint sets of 100 WikiText prompts. This measures one source of fitting variability, rather than providing a complete noise distribution.

Metric, averaged across the band Preview versus 0731 Same-checkpoint, disjoint-corpus reference
Geometry: linear CKA 0.821 0.921
Readout: top-1 token overlap 0.293 0.592
Readout: top-5 token overlap 0.250 0.489
Readout: top-10 token overlap 0.257 0.469

CKA compares the lens maps’ outputs on a fixed input sample. Readout overlap compares the tokens surfaced by each checkpoint’s own lens on the same 20 prompts across 21 layers. Both measurements show lower similarity across releases than in the corresponding same-checkpoint reference.

The geometric difference is greater in the early and middle band: the reference-minus-pair CKA gap averages 0.119 at L19–28 and 0.083 at L29–39. Token-overlap differences are observed across the band without a comparably pronounced early/late contrast. These are useful observations about where the measurements change. However, CKA and token overlap are different metrics; comparing their absolute gaps or ratios does not establish that representational content changed more than geometry.

Behavioral probes and preregistered predictions

Several behavioral measurements retain the same qualitative pattern across the two releases:

Measurement Preview 0731
Next-token agreement under J-ablation, legacy protocol 0.529 0.508
Swap, baseline-conditional 25/52 (0.481) 23/53 (0.434)
Broadcast at α=2 45/192 (0.234) 40/192 (0.208)
Concept readout under suppression 0/13 0/13

This table retains the original measurement protocols for the paired runs. The preview was not included in the corrected 73-configuration sweep, and its agreement should not be silently combined with 0731’s corrected value of 0.523. The table supports similar observed behavior on these probes, without establishing equivalence across checkpoints or invariance to post-training generally.

Five predictions were recorded before the comparison. Their interpretation is:

Prediction What the measurements establish
The band does not move Compatible profiles within L19–39; the preview fit does not independently bracket boundaries outside that window
Geometric divergence concentrates late The measured CKA difference is greater early and mid-band
0731 has sharper ignition Unresolved at the 0.10 interpolation-grid resolution in the sharpest layers
Readout changes more than geometry Both differ relative to their own references; relative magnitude across metrics is not established
Workspace control properties are unchanged No resolved qualitative change on the reported probes; equivalence is not established

The six-set readout profile does change: three scores increase and three decrease. The largest difference is multi-hop intermediate recovery, 0.129 on preview versus 0.294 on 0731. This is a change in what the fitted lens recovers, which can reflect both model representations and the lens’s ability to expose them. It does not directly measure stronger reasoning or establish a causal explanation of the release’s broader capability changes.

The comparison thus shows that the lens is sensitive to differences between these releases, while several coarse behavioral probes remain similar. Identifying which training changes produce those differences requires information about the recipes or more controlled checkpoint comparisons. The comparison scripts and documentation and saved comparison results are available in the public repository.

7. Implications and next experiments

The primary contribution is an executable community replication at a scale where neither the model architecture nor its inference implementation can be assumed to fit the reference code. The rectangular transport and FP8 backward support make the Jacobian lens usable on two substantially different open models, and the experiment suite makes the resulting claims available for further testing.

The results support several components of the proposed workspace account: a middle-layer band and ignition-like readout on DeepSeek, a compatible transition within GLM’s projected window, and strong causal sensitivity to J-lens interventions on both models. They also show why these observations should be evaluated separately from task selectivity. The directions selected by the lens are consequential, yet their ablation changes both multi-hop answers and ordinary next-token predictions.

We did not reproduce the reported reasoning/prediction separation in any of the tested configurations. This limits what the current replication establishes about the global workspace interpretation. A broadly important subspace remains compatible with the causal results, and differences between our implementation and the paper’s remain an alternative explanation for the comparison with Claude.

The main limits concern scope and uncertainty. The large-model grid covers two checkpoints; the smaller-model positive control did not demonstrate detection of selectivity. GLM’s band boundaries remain unmeasured outside the projected window. The normalized estimator and FP8 derivative approximation differ from the reference procedure. WikiText samples overlap the fitting corpus. Most behavioral proportions use 12–192 items, with related prompts and concepts that may reduce the effective sample size; agreement positions are clustered within 20 prompts. Five random seeds quantify control-direction variation, not all of these uncertainties. The broader exploratory study also involves multiple comparisons without a multiplicity correction.

Three follow-ups would be particularly informative. First, establish a successful selectivity positive control and use that same pipeline across capable open models, with independently chosen evaluation items and perturbation magnitudes. Second, fit beyond GLM’s projected window and expand the ignition, arithmetic, and suppression probes to measure their consistency across examples. Third, compare additional post-training releases against multiple disjoint-corpus lens fits, ideally with a verified common pretrained checkpoint.

The public repository contains the port, intervention scripts, evaluation prompts, and selected experimental results; the experiment guide maps the released scripts to the experiments. Fitted lens checkpoints are available on Hugging Face for DeepSeek-V4-Flash-0731, DeepSeek-V4-Flash preview, and GLM-5.2. The 0731 release includes the n=1000 lens, an n=100 fit, and the disjoint-corpus n=100 lens used as the fitting-variability reference. The rectangular DeepSeek lenses require this fork; the square GLM lens can also be loaded with upstream jlens.


Appendix A. Estimator and implementation details

For a prompt with valid positions \(V\), the per-prompt contribution is

\[G_{\ell} = \frac{1}{|V|}\sum_{p\in V} \sum_{\substack{q\in V\\q\ge p}} \frac{\partial\mathbf h_{\mathrm{target},q}} {\partial\mathbf h_{\ell,p}}.\]

The implementation injects a cotangent at all valid target positions simultaneously, then averages the resulting gradients over valid source positions. Causality excludes contributions from targets preceding the source. The target is the final-block residual on the square path and the stream-combination output on DeepSeek. Per-prompt normalization, when enabled, is applied after this position reduction and before averaging across prompts.

The main engineering changes are available in the public implementation:

Component Extension
jlens/protocol.py Optional d_source, d_target, and an out-of-stack target_module, with square-case defaults
jlens/hf.py DeepSeek adapter, four-stream flattening conventions, and explicit control of stream combination before unembedding
jlens/fitting.py, jlens/hooks.py Rectangular derivative accumulation and recording the out-of-stack target
jlens/lens.py Rectangular transport and serialization, including loading older square-lens files
jlens/fp8_autograd.py Activation backward formulas for dense, grouped, and batched FP8 operations
jlens/intervene.py, DeepSeek ablation Swap, steering, and ablation primitives, with a position-adaptive ablation implementation

For GLM’s FP8 matrix multiplication, the activation-gradient formula is \(\nabla_A=\nabla_Y\operatorname{dequant}(B)\), with frozen weight matrix \(B\). It omits the derivative of internal activation quantization. Calling it a surrogate derivative is essential: a straight-through approximation is not the exact derivative of the quantized forward function.

Per-prompt normalization addresses large variation in DeepSeek’s Jacobian gains. On the association readout set, it improves recovery from 0.040 to 0.069. A uniform positive rescaling of an already averaged matrix generally preserves top-token ordering under RMSNorm, apart from numerical effects; changing how prompt matrices are weighted need not. The same distinction applies to normalized token directions used for ablation. The E4 experiment tests the latter change directly at one causal operating point.

In the sweep, directions are formed by transporting vocabulary-centered unembedding rows through \(J_{\ell}\), normalizing them, and constructing a QR basis. The residual’s projection onto that basis is subtracted. A fixed token set and a position-adaptive set therefore test different interventions. On DeepSeek at k=10, their measured norm fractions are approximately 0.0078 and 0.0351, with multi-hop scores of 75/90 and 26/90 respectively.

Swap and steering strength require a separate scale convention. We scale injected unit directions using \(\lVert\mathbf h\rVert_2/\sqrt{d_{\mathrm{source}}}\). Using the full residual norm instead makes the injection 128 times larger for a 16,384-dimensional source. An early run at that scale redirected none of 53 baseline-correct items to the target answer. The failure illustrates why intervention strength must be specified alongside its direction.

Appendix B. Scoring corrections and statistical conventions

The selectivity follow-up separates reproduction of the original runs from evaluation under the corrected protocol. Twenty-nine of the 90 multi-hop prompts end in whitespace. Stripping it changes the unablated scores from 53 to 74 correct on DeepSeek and from 55 to 76 on GLM. The gain is attributable to prompt formatting in this first-token evaluation; the older baseline should not be treated as a pure measure of reasoning capability.

Model Raw-prompt baseline → J-ablation Whitespace-stripped baseline → J-ablation
DeepSeek-V4-Flash-0731 53/90 → 9/90 74/90 → 26/90
GLM-5.2 55/90 → 15/90 76/90 → 29/90

Next-token agreement had a separate inconsistency. The original DeepSeek run scored all 128 positions across 15 WikiText prompts; GLM scored positions 16 through the penultimate position across 20 prompts. Under the consistent latter protocol, DeepSeek’s J-ablation agreement is 0.523, replacing 0.508. GLM’s 0.498 is unchanged. The reproduction run’s random-control agreements are 0.914 and 0.892, replacing the original 0.922 and 0.883. Different random-control runs can also differ because their direction draws differ.

The original “prediction guard” excluded the lens’s top token, despite being described as excluding the model’s top token. The follow-up implements the latter explicitly. DeepSeek’s corrected comparison is 0.523 → 0.594 agreement. Neither guard establishes that all prediction-relevant directions have been protected.

These corrections apply to the experiments actually rerun. The swap results and preview comparison retain their original scoring protocols. A corrected baseline cannot be inserted into a conditional result without rerunning or rescoring the corresponding intervention.

Where given, intervals on item proportions are Wilson 95% intervals. They are descriptive and assume independent items; shared concepts and templates can violate that assumption. Agreement positions cluster within prompts, so position-level binomial intervals are overly optimistic. The selectivity decision is the preregistered threshold conjunction, not a significance test. E2’s five-seed ranges measure randomness in the controls. Neither non-overlapping marginal intervals nor a null difference demonstrates a general mechanism or equivalence between models.

Appendix C. Readout comparison and robustness

The DeepSeek tuned lens learns a 16384 → 4096 map using a next-token KL objective. The logit lens uses the final stream-combination module before normalization and unembedding. The original comparison contains different evaluated layer ranges: the normalized J-lens run searches L0–42, whereas the tuned-lens run searches L19–39. Their results should therefore be read as two within-run comparisons against the corresponding logit baseline, rather than a controlled three-method ranking.

Set J-lens, L0–42 Logit baseline in J run Tuned lens, L19–39 Logit baseline in tuned run
Association 0.069 0.010 0.000 0.010
Typo 0.125 0.052 0.010 0.042
Multi-hop 0.294 0.358 0.233 0.387
Multilingual 0.105 0.213 0.157 0.239
Order of operations 0.145 0.291 0.191 0.291
Poetry 0.031 0.031 0.000 0.061

The readout evaluation and tuned-lens fitting implementations are publicly available. The paired values above keep each evaluation’s own logit baseline; an earlier summary mixed the typo baseline from the tuned run into the J-lens column. The scored J-run item counts are 101/102 for association, 96/96 for typo, 93/93 for multi-hop, 107/107 for multilingual, 55/55 for order of operations, and 98/98 for poetry.

The tuned lens trails its own logit baseline on every set in this experiment. Its low association, typo, and poetry scores are consistent with the paper’s observation that an output-prediction objective can fail to expose unspoken content. The J-lens exceeds its own logit baseline on association and typo, but has clear failures elsewhere. A stricter three-method comparison requires matching evaluated layers, fitting data, and target rules. An absent target in any of these readouts cannot distinguish absent computation from a representation that the vocabulary-based measurement fails to recover.

Fitting corpus size. DeepSeek’s measured n-scaling curve changed little after n=60 on several metrics. A band fit at n=1000 retained the 23/53 conditional swap result observed at n=100. Since the first 100 prompts are contained in the 1000-prompt corpus, an additional lens fitted on a disjoint set of 100 prompts provides a reference for corpus sensitivity. Small readout differences remain, and one such reference does not establish convergence for every measurement.

Frozen attention. At n=100, freezing attention patterns during the derivative calculation leaves association recovery at 0.069; multilingual and order-of-operations scores remain within 12% of the baseline, while multi-hop recovery falls by approximately 41%. This supports robustness for some readouts and a meaningful dependence on attention derivatives for others.

Present-only derivatives. Isolating each source position’s effect on the same target position is more expensive than summing over future targets. Our variant samples one position per prompt. A control samples one source position while retaining future-target terms, separating position sampling from the derivative restriction. For multi-hop recovery, the recorded changes are −0.190 from position sampling and −0.011 from dropping future terms; the latter change is larger on typo recovery, at −0.073. Both variants were also fitted at n=1000, with the reported qualitative conclusions unchanged.

Appendix D. Exploratory observations and measurement lessons

Later arithmetic ranks. In DeepSeek’s arithmetic example, the rank of an intermediate can deteriorate beyond the vocabulary midpoint after another intermediate becomes prominent. This occurs in both preview and 0731. GLM’s corresponding worst ranks remain at or above that midpoint in ranking quality. These single-expression trajectories motivate a broader study of intermediate persistence. A rank comparison against a uniformly sampled vocabulary token does not establish active suppression of the underlying representation.

Base versus post-trained readout. On four self-monitoring probes, three target ranks match between DeepSeek base and 0731, while the roleplay-fictional target moves from rank 24 to rank 0. Reading the post-trained model on raw text preserves the latter result, reducing the prompt-format explanation. However, the released base uses FP8 experts and the post-trained checkpoints use MXFP4, leaving precision confounded with training. Agreement between preview and 0731 does not remove that confound. Four probes are also insufficient to characterize how much of an “assistant perspective” comes from pretraining.

Concept-vector variance. The reported median fraction of concept-vector variance in the selected J-space is approximately 0.4% on 0731 and 0.73% on preview, across 126 concept vectors. GLM’s layer-wise values range from roughly 0.5% to 2.8%, compared with 6–15% reported in the paper. This is a concept-vector decomposition, distinct from the norm fraction of an individual residual edit. The small fractions can coexist with substantial causal effects, but differences in selection and protocol complicate cross-study magnitudes.

Several failed measurements informed the final protocols:

Citation

@article{song2026jacobianlens,
  title={Testing the global workspace at open frontier scale: the Jacobian lens on DeepSeek-V4-Flash and GLM-5.2},
  author={Song, Xiangchen and Feng, Fan and Zheng, Yujia},
  year={2026},
  month={August},
  url={https://xiangchensong.github.io/blog/2026/jacobian-lens-global-workspace/}
}

Revision note (9 September 2026): This revision polishes the exposition, reorganizes the evidence, and clarifies numerical reporting and the scope of the conclusions.

Enjoy Reading This Article?

Here are some more articles you might like to read next:

  • AI and Productive Systems 3: Institutional Design for AI Agents
  • AI and Productive Systems 2: From Tool to Production Subject
  • AI and Productive Systems 1: Intelligence as Infrastructure
  • Understanding Language Models 2: Stable Features and Identifiable Causal Structure
  • Understanding Language Models 1: Mechanistic Interpretability Meets Causal Representation Learning