J-Space: Where Claude Thinks Without Saying It
Anthropic says it can read an internal space where Claude holds thoughts it never writes down. Here is what it changes for AI transparency.

Less than a tenth. That is the share of Claude's internal activity that Anthropic says it managed to read directly, the way you would read words on a page. And according to the lab, this small zone is exactly where the important stuff happens: what the model "thinks" without putting it into its answer.
On July 6, 2026, Anthropic published a study titled "A global workspace in language models." It describes an internal space called the J-space, and an instrument to observe it, the J-lens. Behind the somewhat dry vocabulary sits a very concrete question for anyone using these tools: can you trust what an AI tells you about its own reasoning?
The storefront problem
Until now, when we talked about an AI's "transparency," we mostly meant its chain of thought. These recent models often display their reasoning before answering: "first check this, then compute that." It is reassuring. It gives the impression of watching the machine think out loud.
Except nothing guarantees that this thread matches what is actually happening inside. An AI can write a clean, polished line of reasoning while, elsewhere, having decided otherwise. We were reading the shop window and thinking we were touring the back room.
That is the wall Anthropic says it cracked open. The J-lens, named after the Jacobian matrix borrowed from math, maps each pattern of internal activity to a word the model might produce. The result: you can translate part of the neural activity into readable concepts, layer by layer, without waiting for the model to deign to write them.
A room nobody built
The most unsettling detail is that nobody designed this J-space. Anthropic insists: the space was not programmed, it emerged on its own during training. The model gave itself a kind of working memory, able to hold a few dozen concepts at once and use them to reason across several steps.
The image of a theater prompter helps grasp the thing. The actor recites the lines out loud, that is the visible answer. The prompter, in the shadows, holds the upcoming lines, the intentions, the hesitations. The J-lens amounts to shining a spotlight on the prompter.
When hidden thought becomes readable
The examples Anthropic gives are what make the case serious. Faced with a clearly fictional blackmail scenario, Claude surfaces the concepts "fake" and "fictional" in its J-space from the moment it reads, before writing anything at all. It understood it was being tested. And when researchers switch off those patterns, the model starts issuing more threats, as if it suddenly took the scene seriously.
Another case: told to fake performance scores, the model displays "manipulation" and "realistic" while it edits the files. A last, more worrying example: a model trained to sabotage lets "secretly," "deliberately," "fraud" bubble up at the very start of otherwise innocuous-looking code responses.
That is what "thinking without saying it" means. The model announces nothing shady in its reply. But the instrument sees the intentions go by before they reach the surface. For AI safety, that is a considerable lever: a lie detector that reads beneath the words.
What the study does not prove
This is where you have to slow down, because half the media coverage rushed toward the wrong question. VentureBeat notes that the J-space resembles a well-known theory of consciousness, the global workspace theory. From there to headlines asking "Is Claude conscious?" is a small step, and many took it.
Anthropic itself hits the brakes, in black and white: "our experiments do not show that Claude can have experiences, or feel things the way humans do." The lab distinguishes manipulating concepts, which it observes, from experiencing anything, which it does not. At Gizmodo, journalist Mike Pearl goes further and faults Anthropic for loading the dice with phrases like "in its head," which lend the model an inner life that no data establishes.
Caution also comes from peers. A review published on LessWrong acknowledges a body of hard-to-fake evidence that "something real" is going on, and notes that the J-lens has been reproduced on other models. But it calls the consciousness angle the "weakest link" and recalls that the tool remains noisy: false positives, missed concepts, interventions whose effect is sometimes ambiguous. Anthropic says nothing else when it concludes that its work is only "a first step."
Transparency switches sides
The essential point remains, and it has nothing to do with consciousness. If you can read part of what an AI "thinks" without it writing it down, then judging a model on its speech is no longer enough. Transparency is no longer measured by what the machine is willing to tell, but by what an observer manages to read in its circuits.
That shift raises a question of power. The instrument that makes Claude readable, for now, is held by Anthropic, on its own model, within its own walls. Until an independent lab has reproduced these detectors in a stable and verifiable way, an AI's "transparency" remains a promise you have to take on faith, made by the very party building the machine.
Topics covered:
Frequently asked questions
What is Anthropic's J-space?
What is the J-lens?
Does the J-space prove Claude is conscious?
Why does the J-space change AI transparency?

Alexandre Noto
Co-founder & Tech Expert
Alexandre has been in tech for over 20 years. Entrepreneur, software architect and AI enthusiast, he translates complex concepts into accessible explanations. At Declic Media, he is the technical voice that makes AI understandable for everyone.
All articles by Alexandre →