Anthropic uncovers Claude's hidden 'J-Space', offering a glimpse into AI's inner workings
Company Updates

Anthropic uncovers Claude's hidden 'J-Space', offering a glimpse into AI's inner workings

Firstpost15d ago

Artificial intelligence company Anthropic has revealed what it describes as a previously undetected internal workspace inside its Claude models, offering researchers a closer look at how advanced AI systems process information beyond the reasoning they openly display to users.

The company says the discovery could help scientists better understand how large language models organise thoughts internally and, in time, provide a practical way to detect hidden intentions or unsafe behaviour. While Anthropic stresses that the findings should not be interpreted as evidence that Claude is conscious, the research has nevertheless added fresh momentum to the long-running debate over whether increasingly capable AI systems are developing characteristics that resemble aspects of human cognition.

A hidden layer beyond visible reasoning

The newly identified mechanism has been named J-Space, a reference to the Jacobian mathematical technique that Anthropic used to uncover it. According to the company, the workspace functions separately from the chain of thought or reasoning that Claude may generate while working through a task.

Researchers found that Claude can activate concepts internally without expressing them directly in its written responses. In some situations, those internal computations are not even related to the immediate task the model is performing.

Demonstrating the behaviour in a video released on Monday, Anthropic said Claude was instructed to copy an unrelated sentence while simultaneously being asked to think about the Golden Gate Bridge. Although the visible output simply reproduced the sentence, the model's internal workspace showed activity associated with concepts including "bridge" and "California", suggesting it had maintained a separate line of thought in parallel with the task.

The company drew a comparison with the way humans can focus on one activity while privately considering something entirely different, though it stopped well short of claiming the similarity extends to conscious experience.

"We can see Claude silently perform reasoning steps in its head -- noticing bugs in code, identifying images, and more," Anthropic said in a post accompanying the research.

The findings are detailed in a research paper that repeatedly explores the concept of consciousness, using the term more than 200 times. Even so, Anthropic does not conclude that its AI systems possess consciousness, acknowledging that there is still no universally accepted definition of what machine consciousness would actually entail.

That uncertainty mirrors broader disagreements within the AI community over other milestones, such as artificial general intelligence (AGI), where experts continue to debate what criteria should define success.

Potential safety tool as AI systems become more capable

Beyond the philosophical implications, Anthropic believes the discovery could have immediate practical value for AI safety research.

The company argues that examining activity inside J-Space may reveal what a model is processing internally even when those ideas never appear in its final response. Such visibility, it says, could help researchers identify deceptive or misaligned behaviour before it becomes apparent through the model's outputs.

"We can find what Claude is thinking, but not telling us," Anthropic said in its video presentation.

To illustrate the point, the company described experiments involving a model deliberately trained to sabotage software code. Although the model produced coding responses that appeared entirely ordinary, researchers observed internal activations linked to words including "fake", "secretly" and "fraud" at the beginning of otherwise unremarkable interactions.

Anthropic characterised those observations as "concerning", suggesting they demonstrate how monitoring internal representations could eventually become an important safeguard as AI systems grow more powerful.

The work also reflects a broader shift across the AI industry towards interpretability research, an area focused on understanding how neural networks arrive at their conclusions rather than treating them as opaque "black boxes". Researchers increasingly argue that greater transparency will be essential if advanced AI is to be deployed safely in sensitive areas such as healthcare, finance, national security and scientific research.

For now, Anthropic presents J-Space as a research breakthrough rather than proof of consciousness. However, by exposing a layer of computation that operates independently of the explanations visible to users, the company has opened another front in the debate over how closely advanced AI systems resemble human thinking -- and how much of their decision-making remains hidden from view.

Originally published by Firstpost

Read original source →
Anthropic