Anthropic describes J-space hidden tokens in AI models

Anthropic says it identified “J-space,” a latent structure of non-output tokens that can steer model reasoning, though reliability remains unclear.

Jason Kwon ·

Anthropic describes J-space hidden tokens in AI models

Anthropic researchers said they have identified a latent internal structure inside large language models that they call “J-space,” describing it as a region of non-output tokens that can influence how systems tackle complex tasks.

The company said the finding offers a new way to examine the multi-billion-parameter computations that drive model behavior, by surfacing internal elements that do not appear in the final response but still shape how answers are produced.

Hidden “words” that do not appear in outputs Hidden “words” that do not appear in outputs Anthropic said the newly described space contains internal Anthropic said the newly described space contains internal “words” that never show up in model outputs, yet function as structural guides during computation. Researchers reported that these latent tokens can indicate progress through a task or switch on specific behavioral pathways. According to the company, some of these hidden words resemble momentary signals of recognition or internal commentary about ongoing operations. The researchers presented the work as a practical handle for inspecting what would otherwise be an enormous set of interacting calculations. A probing technique and evidence of active use Engineers at Anthropic said they developed a specialized technique to inspect the model and expose these latent pathways. Using that method, they reported evidence that the system can identify and manipulate words within this internal space.

Anthropic said this supports the view that the Anthropic said this supports the view that the model actively uses the hidden architecture during decision-making, rather than the tokens being merely a passive artifact of training.

Claude example: an internal “panic” signal As an illustration, the company described a case in which the Claude model changed its strategy on a coding test when the word “panic” appeared internally. Anthropic framed the example as a sign that internal tokens can coincide with shifts in how the system proceeds through a task.

Mechanistic interpretability, control, and remaining uncertainty Anthropic Chief Executive Officer Dario Amodei has previously said that understanding internal model mechanics is necessary for full control of these systems. The company described its work as progress in mechanistic interpretability , a field aimed at decoding the contributing signals behind computational outcomes.

In Anthropic’s description, J-space adds another lens for examining internal operations that are not directly visible through normal inputs and outputs. At the same time, the material presented said it remains an open question how reliably these latent tokens can explain model behavior across different tasks.

Limits of biological analogies and governance relevance

Anthropic noted that researchers sometimes rely on biological analogies when discussing internal processes, while observers have cautioned against anthropomorphizing language models. The company emphasized that these systems operate through mathematical relationships that trigger cascades of calculations rather than biological cognition.

As institutions increase their reliance on large language models, Anthropic said the ability to audit internal reasoning pathways is expected to influence how artificial intelligence governance frameworks are designed. The company positioned J-space as a potential tool for that kind of auditing, while underscoring that model operations remain fundamentally mathematical.

More stories