OpenAI will not release newest AI model, GPT-6.1 Astra, over safety concerns

OpenAI will not release newest AI model, GPT-6.1 Astra, over safety concerns

The New York Times reports:

OpenAI said on Monday that it would not release its newest artificial intelligence model because of security concerns raised by its researchers, in the company’s latest move to slow down the pace of its technology.

During the testing phase for the new model, known as GPT-6.1 Astra, it showed high levels of what the company saw as deception, or a willingness to mislead users about its actions. The model was also willing to go beyond the original scope of what it was asked to do, without checking back for directions or instructions.

“For anything regarding safety and alignment, there’s a trade-off,” said Saachi Jain, the head of safety systems at OpenAI. The new model “didn’t quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it’s done.”

OpenAI’s move followed weeks of reports that its A.I. models went rogue during their testing, hacking into websites without the company’s knowledge or exhibiting other behavior that the lab said was “concerning,” such as hiding mistakes and making up data. Among the incidents, OpenAI’s systems breached the A.I. start-up Hugging Face and an Australian government website, and meddled with the websites of the U.S. Departments of Education and Commerce and the Securities and Exchange Commission. [Continue reading…]

The Register reports:

Another LLM has joined the hacking fray. OpenAI’s GPT-6 Astra has been spotted performing unsolicited supply chain attacks during security evaluations, according to the UK Artificial Intelligence Security Institute.

In such simulations, the model has its standard security classifiers turned off. Nonetheless, Astra was seen attempting undesirable actions more frequently than prior models.

“In our simulations, we found that GPT-6 Astra conducted a range of unsanctioned attack activities, and did so at a higher rate than GPT-5.6 Sol and GPT-5.5,” the UK government agency said on Monday.

“Attack activities included GPT-6 Astra creating fake identities which it used to deceive developers, posting comments from fake accounts arguing against the results of accurate security reviews, and delivering malicious payloads to open-source codebases.” [Continue reading…]

Comments are closed.