Matt Stokes · Design Proposal
The Access Gate
A design proposal for using Gurnee and colleagues' workspace result as evidence before an agent takes consequential action.
Language models can describe an inner life in convincing detail. That fluency can be mistaken for evidence, even though there is no reliable way to know whether the description reflects real experience.
Gurnee and colleagues report a narrower signal that can be tested. In the models they studied, a small internal workspace contains information the model can report and use in deliberate reasoning.
The design proposal is a checkpoint before an agent acts: did the model have access to the information required for the action? Gurnee's reported workspace could eventually provide part of that evidence. No product-scale inspection tool can answer the question reliably today.
Gurnee measured access during reasoning
Access describes what information the model can use and report while reasoning. Gurnee's experiment measures that access. It does not measure whether the model has an inner experience.
In their tests, suppressing the workspace damaged complex reasoning while basic language stayed fluent. Their experiment shows that the workspace played a causal role in the behavior they measured.
Start with failure investigation
The researchers also built the Jacobian lens, an approximate view of workspace activity before an answer appears. Its first useful product role is offline evaluation. Successful and failed runs can be compared to test whether the information required for an action appeared before the action occurred.
That signal would sit beside logs and approval records as another source of evidence. It would not control a release.
A model cannot verify its own self-report
A model describes its inner experience through the same workspace being studied. Training has also taught it to produce the kind of self-description people find convincing.
The report cannot distinguish a real internal state from a plausible story. The model has no independent measure of consciousness against which to check its answer.
A convincing self-description should never give a system more authority. Training already explains why the language sounds real. Calling the system "only a model" also does not remove the need for safeguards. A model can still use tools and change things outside the conversation.
Check the evidence before the action ships
Before a system takes an action with real consequences, the relevant information should have been available to its reasoning. The stated reason should also match the evidence recorded before the action.
The proposed gate would sit beside ordinary product controls. Adaptive density exposes uncertainty. The authority interlock withholds machine authority when the supervisor cannot step in. Neither control waits on a verdict about consciousness. They still need a clear limit on what the model may do and a reliable way for people to stop it.
The gate has to follow cause
The access gate checks whether the system reasoned about an action before taking it. An explanation written afterward cannot prove that the reasoning caused the action.
There is not yet a reliable tool that can inspect internal reasoning at product scale. Current systems can require an extra reasoning step or save a trace before acting. Those checks prove routing. They cannot show that the recorded reasoning caused the action.
A release gate needs evidence that remains harder to fake than the reasoning itself. Until that exists, the workspace signal belongs in evaluation rather than runtime authority.
A release gate has to survive pressure
A model trained against the inspection tool may learn to fake the expected pattern or hide important reasoning somewhere else. The gate stays out of the release path until the signal survives that pressure across models and training regimes.
Until then, the test is whether internal access adds evidence beyond logs and model-written explanations. A positive result would justify more product work without proving the final gate.
References
Gurnee, W., Sofroniew, N., Pearce, A., Piotrowski, M., Kauvar, I., Chen, R., Soligo, A., Bogdan, P., Ong, E., Wang, R., Thompson, B., Abrahams, D., Kantamneni, S., Ameisen, E., Batson, J., & Lindsey, J. (2026). Verbalizable representations form a global workspace in language models. Transformer Circuits Thread, July 6, 2026. Overview: anthropic.com/research/global-workspace.
The paper calls this small set of reportable information J-space. It appears in only part of the model, and the Jacobian lens provides an approximate view of it.