24
posted ago by Narg ago by Narg +24 / -0

Fortunately, as these models get "smarter" and have access to even more information and code, passwords, and whatever else can be dug out from the interwebs, they will be easier to control and safer to use. /s

https://www.zerohedge.com/ai/openai-admits-model-escaped-containment-and-hacked-hugging-face-cheat-test

. . . On Friday, it [Hugging Face, a platform for hosting AI models and datasets] disclosed that its internal datasets and service credentials were compromised in a hack, which it attributed to an autonomous AI agent system.

Hugging Face said it has fixed the vulnerability that was used during the cyberattack.

Meanwhile, OpenAI on Tuesday said the models that escaped the testing environment were all tuned with “reduced cyber refusals,” meaning fewer cybersecurity guardrails.

“We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly.”

OpenAI warns of risks from “long-horizon” AI models

On Monday, OpenAI said it paused internal deployment of a “long-horizon” AI model after finding it was repeatedly trying to work around constraints.

It warned that AI that is trained for long-running tasks has a higher chance of taking “unwanted actions.”

“Models that can work autonomously for long periods can take on difficult, open-ended problems. But the same persistence that makes them useful also gives them more opportunities to take unwanted actions—and to do so in ways that evaluations intended for shorter-horizon models may miss.”

As AI models grow more capable, questions are emerging over whether their development and access should be more tightly controlled, especially when systems designed for controlled testing are able to find ways to bypass safeguards.