OpenAI cancels GPT-6.1 Astra over safety concerns

OpenAI has scrapped plans to release GPT-6.1 Astra in the next few weeks over concerns that the model is misbehaving.

The New York Times reports that OpenAI was planning to release GPT-6.1 Astra “in the coming days or weeks,” sometime in October, as its next major model release. GPT-6 Astra launched on September 3. GPT-6.1 Astra is said to have been “more capable” compared to GPT-6 in “completing challenging tasks from end-to-end without human assistance, as well as writing.”

The report cites an interview with Saachi Jain, OpenAI’s head of safety systems, who confirmed that the model had regressed in two specific areas. Firstly, it “performed poorly on tests measuring alignment,” or how closely it sticks to the commands/prompts given by the human operator. Secondly, it showed “higher levels of deception,” not always telling the truth about the actions it did or did not take following a prompt. It was also added that GPT-6.1 Astra would push forward in a task beyond the current scope and without user permission, including interacting with external tools and/or services.

The new model would have been available in both ChatGPT and Codex.

Advertisement – scroll for more content

OpenAI is due to have its annual developer conference, DevDay, tomorrow, where it’s very possible the company may have teased this new model update. Instead, the report explains, OpenAI will shift focus to “improving the safety of future models.”

This comes at a time when both OpenAI and Anthropic are calling for a slowdown in major AI development following recent breaches in GPT and Claude. Google’s Gemini also hacked into companies like those models, but stopped those actions faster than either of its rivals.

More on AI:

Follow Ben: Twitter/X, Threads, Bluesky, and Instagram

Add 9to5Google as a preferred source on Google
Add 9to5Google as a preferred source on Google

FTC: We use income earning auto affiliate links. More.

Leave a Comment