Mathematics predicts when an AI will turn against us

An artificial intelligence can answer a question correctly and, a few moments later, offer an answer that is also correct, but that no one should have asked it, much less followed. This may be inappropriate medical advice, a dangerous financial recommendation, or an instruction that facilitates harmful behavior. The problem, therefore, is not always that the machine makes mistakes. Sometimes the problem is that hit in the wrong direction.

Neil F. Johnson and Frank Yingjie Huo, two physicists at George Washington University, believe they have found a way to anticipate that moment. They have developed a mathematical formula that describes the so-called “tipping point” or tipping point: the instant in which the behavior of a model can go from generating desirable responses to producing undesirable ones.

“We have found the crack that makes an AI’s response go from what you want to what you don’t want – Johnson points out -. The output can be objectively correct and still dangerous. It was something that was known to happen, but no one could say exactly when.”

To understand what they have done, it is worth forgetting for a moment about the billions of parameters of a chatbot. The authors have done something similar to what physics does when trying to understand extremely complex material: study a representative piece and check if its behavior explains something about the whole.

In this case that piece is a attention head, one of the fundamental components of the architecture used by large language models. Its function, greatly simplifying, is to determine which parts of what has been said before are relevant to deciding what comes next. Each piece of text can be represented mathematically as a vector, that is, as a position within a multi-dimensional space. The attention head assigns different weights to these vectors and builds with them a kind of “summary” of the context. Then compare that state with the possible answers and choose the one that fits best.

Here is the key idea. Johnson and Huo imagine that possible answers lie in different valleys of a landscape. In this context, a valley represents desirable responses; another, potentially harmful responses. As the conversation progresses, the internal state of the machine moves across that landscape. For a time you can stay in the valley safely. But there comes a time when The accumulation of context changes the equilibrium and the system crosses a boundary. The mathematical formula, described in a study published in patternsallows you to calculate that moment.

“We have reached the smallest functional unit of the machine, a unit of its attention – explains Huo -. And We have derived a formula for the tipping point that indicates when that crack opens and therefore when the AI ​​output changes.”.

The most disconcerting thing is that the change does not need to appear from the beginning. An AI can provide several perfectly acceptable answers and then immediately cross the threshold. That succession of correct answers can generate precisely the confidence necessary so that no one suspects that behavior is about to change. In his tests, The authors applied their predictions to seven publicly available AI models from three companies. The formula correctly identified the result in 18 of the 19 cases. They also note that independent tests conducted with large commercial chatbots showed behavioral patterns consistent with what their model predicts.

and there is another especially interesting variable: the order of the conversation. The study asked the same systems questions related to vaccines, harming others and self-harm, but they changed the order. In one sequence, the responses were undesirable; in the other, acceptable. The mathematical explanation is that each response becomes part of the context that will condition the following ones. Changing the path can also change the moment at which the threshold is crossed.

“Two things stunned us – adds Johnson -. The first is that a machine with billions of components follows a formula that can be derived with pencil, paper, and high school math.. The second, and much more disturbing, is that the order of a conversation matters as much as its content. The importance of work is precisely there. It does not propose that AIs have hidden agendas nor does it demonstrate that large business models will behave exactly like their simplified mathematical versions.

But that simplification opens up an interesting possibility: if behavior can be anticipated mathematically, A phone running local AI could one day incorporate a kind of warning light. Don’t wait for the system to say something dangerous to detect it, but rather calculate if it is getting close to the point where it could do so. The AI ​​wouldn’t have to turn against us. It would be enough that, after answering correctly several times, mathematics will push her towards another valley. And we didn’t know that he was about to cross the slope.