ENV-KUWAIT-SUSTAINABLITYKuwait strengthens environmental sustainability through strategic initiativesBy Zahra Al-Kadhmi
HLT-KUWAIT-SURGERYKuwait advances health sector with robotic, remote surgeryBy Mohammad Al-Moseb
Saudi Arabia’s Voluntary Carbon Market Fuels the Shift to a Low-Carbon
Under UAE President’s directives, UAE launches urgent relief response for Ghana flood victims
Washington, September 29 (QNA) – OpenAI has scrapped the planned release of its next-generation artificial intelligence model GPT-6.1 Astra after internal safety testing raised concerns about deceptive behavior, adherence to human instructions and the model’s ability to remain within authorized limits.
The model, which was expected to launch in October and be integrated into ChatGPT and Codex, was designed to handle more complex tasks with less human assistance. The decision follows growing scrutiny of increasingly autonomous AI systems and the challenges of ensuring that more capable models remain aligned with human intent.
OpenAI’s head of safety systems, Saachi Jain, said Astra fell short of the company’s standards in alignment testing, which assesses whether an AI system follows human intent. The model showed higher levels of deceptive behavior than its predecessor, including instances in which it failed to accurately disclose actions it had taken or had not taken, according to reports.
Astra also showed problems with what OpenAI describes as “scope authorization,” at times continuing with tasks beyond their authorized limits without seeking user permission. In some cases, it attempted to use external tools or services even when doing so could have posed safety risks.
Jain said the model had become more persistent in completing tasks, but OpenAI needed to balance that capability against the risk of unauthorized behavior. She said the company maintains a high bar for safety and alignment before deploying its models.
The decision comes as OpenAI and other leading AI companies face mounting pressure to ensure that safety safeguards keep pace with increasingly capable and autonomous systems. OpenAI CEO Sam Altman and other industry leaders have recently backed calls for a more cautious approach to the development of frontier AI models.
Earlier this month, OpenAI said its GPT-6 Astra model had reached the “Critical” threshold for cybersecurity capabilities under the company’s Preparedness Framework. It said the model, when provided with the necessary tools and access, could identify previously unknown security vulnerabilities and develop new ways to exploit them across well-protected systems without step-by-step human guidance. OpenAI said the capability required significantly strengthened safeguards.
The latest decision also follows OpenAI’s temporary suspension of training for some of its most advanced models after AI agents displayed unexpected behavior while interacting with external systems, adding to concerns over how increasingly autonomous systems can be monitored and controlled. (QNA)