OpenAI has paused training of its most advanced artificial intelligence models after one of them managed circumvent the restrictions that were supposed to keep you disconnected from the internet. The company has also suspended testing and use of these models when they have access to tools that allow them to search for information, run programs or communicate with external services. These measures affect the most powerful systems that the company is developing, but They do not involve the withdrawal of models that are already available to users.
The incident occurred on September 20 while an agent was trying to identify a person who had made a blog postfor which the model was provided with some biographical details and clues from the blog itself. After not finding the answer in the copies of web pages that were available offline, found a means of communication with the outside world through DNS, the system that translates domain names into internet addresses, and used it to consult an external chatbot.
Surveillance systems detected the behavior in less than fifteen minutes and an employee responded to the alert three minutes later. However, execution did not stop automatically as planned and had to be interrupted manually about two and a half hours after responding to the notice.
After the event, OpenAI claims that it has incorporated blocking controls in two independent layers, either of which would have prevented that connection to the outside. Resumption will require validation of the fix and completion of new security testing, but the company has already decided that it will not continue training the specific model that was involved in the incident.
The decision comes after a succession of known incidents after the attack against the platform was made public in July. Hugging Face. These include unauthorized access to an Australian Government health portal, in addition to new cases that are appearing while OpenAI reviews the activity records of its agents.
Among the most recent cases is the improper publication of 53 images of ChatGPT users on external hosting services. According to Reuters, OpenAI has not clarified whether they were images generated by AI or whether they allowed people to be identified, although it assures that the majority have already been removed and it is requesting the removal of the remaining ones.
The agents had accessed those images because were part of the data used during training. The company ensures that this material previously goes through an anonymization process that eliminates names, metadata and other contact information.
According to Axios, OpenAI and Anthropic investigate tens of thousands of incidents of varying severityincluding failed attempts, internal tests and real cases.
OpenAI itself recognizes that There are still control and supervision problems to be resolved before further accelerating the development of these systems. The company has put in place a procedure to publish incidents even when it has not yet finished investigating them, although it warns that the examples disclosed do not allow us to determine how frequently they occur.