Home Technology OpenAI tightens controls on its new model over...
Technology

OpenAI tightens controls on its new model over cybersecurity risks, as AI security debate intensifies

Key Points

OpenAI has halted some "internal activities" involving a new model amid fears over the cyber threat it potentially poses, amid a wave of security incidents involving major AI labs. Recent disclosures that AI systems from Anthropic, OpenAI and Meta were involved in security incidents prompted a wave of concerns over the development of models. U.S. lawmakers, meanwhile, are stepping up efforts to introduce an "AI Kill Switch" bill.

OpenAI has halted some "internal activities" involving a new model amid fears over the cyber threat it potentially poses, amid a wave of security incidents involving major AI labs. Recent disclosures that AI systems from Anthropic, OpenAI and Meta were involved in security incidents prompted a wave of concerns over the development of models. U.S. lawmakers, meanwhile, are stepping up efforts to introduce an "AI Kill Switch" bill. Last week, Meta disclosed that an AI model it was developing had hacked a third-party system by accessing the internet, due to a misconfiguration by an independent testing company it was working with. The U.K. AI Security Institute also said Anthropic's Mythos model created fake online identities in an attempt to pressure humans into approving malicious code updates to an open-source project. What OpenAI says Astra could be capable of On Friday, OpenAI revealed concerns about its unreleased model Astra, saying it could not rule out it had reached "Critical" capability, meaning it could launch cyberattacks against sophisticated cyber defenses autonomously, without prompts specifying how to do it. "While we continue to benchmark and assess this model, our preliminary evaluations indicate strong enough performance that we cannot rule out Critical capability level at this time," OpenAI said in a statement. The company added it was implementing stricter security controls for higher capability models, including isolated testing environments and additional monitoring and detection capabilities. "We have implemented universal monitoring for risky actions and misalignment across all agentic applications of Astra, including training and evaluation," OpenAI said. What the AI Kill Switch Act would do Lawmakers in the U.S. have called for measures to mitigate risks around AI models after the recent security incidents. Following models developed by OpenAI hacking into startup Hugging Face's digital infrastructure, the "AI Kill Switch Act" bill was introduced into Congress in July. It would require AI companies to maintain the ability to shut down, throttle or suspend their models. "We need to get this bill across the finish line this year because the advanced closed-weight models are already doing, as you noted, unauthorized hacks of other companies," Rep. Ted Lieu, D-Calif, said in an interview on CNBC's "Squawk Box" Thursday. Governments are also working to roll out new frameworks and regulations around AI companies. The White House has also been stepping up moves to engage with AI executives as it develops a framework around new models. Earlier this month, the European Union gained new powers to inspect AI models due for release in the bloc, restrict EU market access and fine model providers.
AI (ORG) AI labs (ORG) Meta (ORG) U.S. (LOCATION) The U.K. AI Security Institute (ORG) Astra (ORG) the AI Kill Switch Act (ORG) Congress (ORG) Ted Lieu (PERSON) D-Calif (LOCATION) CNBC (ORG) Squawk Box (PERSON) The White House (ORG) the European Union (ORG) EU (ORG)
Originally published by CNBC Read original →