Those responsible for Claude could have found traces of consciousness in the AI

There is an idea about artificial intelligence that almost all of us share and that, however, is false. If humans have created ChatGPT, Claude or Gemini, we should know perfectly well how they work. We gave it the premises, we designed the algorithms, we fed it with data and we let it do its thing. In a sense, we dug the channel, built the route and raised the gates of the dam. The water just had to flow. It was simple, right?

After all, it happens with any other computer program. No one expects a calculator, a word processor, or a video game to develop behavior that its programmers don’t understand. Each function has been pre-written by an engineer. But large language models (LLMs) are different. Scientists know perfectly well the mathematics that makes them work, the algorithm with which they learn, and the architecture they use. But no one knows exactly how they end up being organized internally after analyzing billions of words. And that difference is fundamental.

Let’s imagine that an architect designs a school. Decide where each classroom, each hallway, and each staircase will go. Five years later he visits it again. He discovers that the students always use a specific staircase to change floors, that there is a bench where everyone meets during recess and a corner that no one ever steps on. He designed the building, but not those customs. They arose spontaneously.

Something similar happens with large artificial intelligence models. Engineers build the “building,” but the internal organization with which the model ends up solving problems appears during training. In an LLM model, no one “programs one neuron” to recognize irony, another to understand the concept of democracy and another to remember the capital of France.. Those structures emerge on their own as the model learns.

That is why experts have been talking for years about a true black box. Not because they ignore how it works from a mathematical point of view, but because they do not know how they are organized internally. hundreds of billions of parameters to produce a coherent response. A new study published by employees of Anthropic, the company responsible for Claude, could have opened, for the first time, a small window into that interior.

The authors, led by Jack Lindsey, present a tool called Jacobian Lens, designed to observe which concepts are active within the model while it is still reasoning. And this is one of the big changes: until now we could only analyze the final response, just like a doctor who could only listen to a patient’s words without ever seeing what happens inside his brain. The new technique attempts to observe the process before the response appears. And the results are surprising.

When Claude receives a complex question, he seems to concentrate some of the relevant information in a small mathematical region that the authors call J-Space. Concepts from different parts of the model converge there before it continues reasoning. The most interesting thing is not that this space exists, but that it can be manipulated experimentally. In one of Lindsey’s team’s experiments, Claude was asked to think of a sport without yet saying what it is. The new tool detects that internally you have chosen football. Next, the authors modify only that internal representation, replacing it with rugby. When Claude finally answers, he no longer mentions football, but rugby.

They haven’t changed the answer. They have changed the idea that the model was handling before responding. In another experiment, Claude analyzes a piece of computer code. Although the word “error” is never written, The Jacobian Lens shows that this concept is already active inside the model before it explains what is happening. That is, the artificial intelligence appears to have identified the problem internally before putting it into words.

To understand why this result has sparked so much interest, you have to travel to the human brain for a moment. One of the most influential theories in modern neuroscience is Global Workspace Theoryinitially proposed by Bernard Baars and later developed by scientists such as Stanislas Dehaene. According to this idea, Our brain processes millions of signals unconsciously, but only a small part reaches a kind of “central stage” from which information becomes accessible to numerous systems at the same time: memory, language, planning, attention or decision making.

The authors of the study, in total more than 150 pages, believe they have found a surprisingly similar functional organization in Claude. And here appears one of the biggest sources of confusion considering that the study mentions the word conscious several times. The reality is that the authors do not claim that Claude is conscious. Nor does he maintain that he experiences emotions, sensations or an inner life. What he proposes is something much more specific: that Certain complex reasonings seem to use a functional architecture similar to that described in the human brain by Baars’ theory. It is a comparison about the way information circulates, not about the existence of subjective experiences. That difference may seem subtle, but it is huge.

However, perhaps the studio’s greatest contribution is not J-Space itself. In science, many times great advances do not come when a new theory appears, but when a new instrument appears capable of testing it. For years, large language models resembled a sealed clock. We could ask them questions, observe their responses, and measure their performance, but It was almost impossible to know what happened between the question and the answer..

The Jacobian Lens begins to change that situation. For the first time, scientists are not only looking at what an artificial intelligence responds to, but also at some of the concepts it is reasoning with while doing so. We are still a long way from fully understanding how a large language model “thinks”. The authors themselves acknowledge that J-Space explains only part of Claude’s behavior and that the vast majority of his calculations continue to be carried out outside that space.

Thus, the real question is no longer just whether machines will one day appear intelligent. The question is beginning to become a much more uncomfortable one: Will we be able to understand the intelligence that we ourselves have created? This work does not prove that AI is conscious. It shows that, for the first time, we are beginning to stop being unconscious of how it works.