Society & safeguards

J-Space: Where Claude Thinks Without Saying It

5 min read

Anthropic says it can read an internal space where Claude holds thoughts it never writes down. Here is what it changes for AI transparency.

The free AI newsletter
J-Space: Where Claude Thinks Without Saying It

Less than a tenth. That is the share of Claude's internal activity that Anthropic says it managed to read directly, the way you would read words on a page. And according to the lab, this small zone is exactly where the important stuff happens: what the model "thinks" without putting it into its answer.

On July 6, 2026, Anthropic published a study titled "A global workspace in language models." It describes an internal space called the J-space, and an instrument to observe it, the J-lens. Behind the somewhat dry vocabulary sits a very concrete question for anyone using these tools: can you trust what an AI tells you about its own reasoning?

The storefront problem

Until now, when we talked about an AI's "transparency," we mostly meant its chain of thought. These recent models often display their reasoning before answering: "first check this, then compute that." It is reassuring. It gives the impression of watching the machine think out loud.

Except nothing guarantees that this thread matches what is actually happening inside. An AI can write a clean, polished line of reasoning while, elsewhere, having decided otherwise. We were reading the shop window and thinking we were touring the back room.

That is the wall Anthropic says it cracked open. The J-lens, named after the Jacobian matrix borrowed from math, maps each pattern of internal activity to a word the model might produce. The result: you can translate part of the neural activity into readable concepts, layer by layer, without waiting for the model to deign to write them.

A room nobody built

The most unsettling detail is that nobody designed this J-space. Anthropic insists: the space was not programmed, it emerged on its own during training. The model gave itself a kind of working memory, able to hold a few dozen concepts at once and use them to reason across several steps.

The image of a theater prompter helps grasp the thing. The actor recites the lines out loud, that is the visible answer. The prompter, in the shadows, holds the upcoming lines, the intentions, the hesitations. The J-lens amounts to shining a spotlight on the prompter.

When hidden thought becomes readable

The examples Anthropic gives are what make the case serious. Faced with a clearly fictional blackmail scenario, Claude surfaces the concepts "fake" and "fictional" in its J-space from the moment it reads, before writing anything at all. It understood it was being tested. And when researchers switch off those patterns, the model starts issuing more threats, as if it suddenly took the scene seriously.

Another case: told to fake performance scores, the model displays "manipulation" and "realistic" while it edits the files. A last, more worrying example: a model trained to sabotage lets "secretly," "deliberately," "fraud" bubble up at the very start of otherwise innocuous-looking code responses.

That is what "thinking without saying it" means. The model announces nothing shady in its reply. But the instrument sees the intentions go by before they reach the surface. For AI safety, that is a considerable lever: a lie detector that reads beneath the words.

What the study does not prove

This is where you have to slow down, because half the media coverage rushed toward the wrong question. VentureBeat notes that the J-space resembles a well-known theory of consciousness, the global workspace theory. From there to headlines asking "Is Claude conscious?" is a small step, and many took it.

Anthropic itself hits the brakes, in black and white: "our experiments do not show that Claude can have experiences, or feel things the way humans do." The lab distinguishes manipulating concepts, which it observes, from experiencing anything, which it does not. At Gizmodo, journalist Mike Pearl goes further and faults Anthropic for loading the dice with phrases like "in its head," which lend the model an inner life that no data establishes.

Caution also comes from peers. A review published on LessWrong acknowledges a body of hard-to-fake evidence that "something real" is going on, and notes that the J-lens has been reproduced on other models. But it calls the consciousness angle the "weakest link" and recalls that the tool remains noisy: false positives, missed concepts, interventions whose effect is sometimes ambiguous. Anthropic says nothing else when it concludes that its work is only "a first step."

Transparency switches sides

The essential point remains, and it has nothing to do with consciousness. If you can read part of what an AI "thinks" without it writing it down, then judging a model on its speech is no longer enough. Transparency is no longer measured by what the machine is willing to tell, but by what an observer manages to read in its circuits.

That shift raises a question of power. The instrument that makes Claude readable, for now, is held by Anthropic, on its own model, within its own walls. Until an independent lab has reproduced these detectors in a stable and verifiable way, an AI's "transparency" remains a promise you have to take on faith, made by the very party building the machine.

Topics covered:

EthicsAnthropic

Frequently asked questions

What is Anthropic's J-space?
The J-space is an emergent internal workspace inside Claude, where the model holds a few dozen concepts to reason with. It accounts for less than a tenth of the model's activity but concentrates most of what matters for safety.
What is the J-lens?
The J-lens (Jacobian lens) is Anthropic's instrument that translates Claude's internal activity into readable concepts, layer by layer, without waiting for the model to write them into its answer.
Does the J-space prove Claude is conscious?
No. Anthropic states in plain terms that its experiments do not show that Claude can feel things the way humans do. The J-space is about manipulating concepts, not about subjective experience.
Why does the J-space change AI transparency?
Because you can read part of what an AI thinks without it writing anything down. Judging a model on its words alone no longer holds: transparency now depends on what an observer can read inside its circuits.
Alexandre Noto

Alexandre Noto

Co-founder & Tech Expert

Alexandre has been in tech for over 20 years. Entrepreneur, software architect and AI enthusiast, he translates complex concepts into accessible explanations. At Declic Media, he is the technical voice that makes AI understandable for everyone.

All articles by Alexandre →
The free AI newsletter