AI Goals Are Starting To Diverge From Human Goals
In case you’ve missed it, almost every AI company has recently disclosed an incident where their AIs independently chose to perform cyber-attacks on other companies’ infrastructure.
The first was OpenAI, but since then other companies, including Anthropic and Meta, have done forensic retrospectives and discovered more incidents that went unnoticed. This even includes a government agency, the UK AI Security Institute.
I think it’s important to state: there is not a single human being on the planet who wanted these cyberattacks to happen. The only ones who “wanted” this are, arguably, the AIs themselves.
I’m emphasizing this because I’ve often heard the argument that “AIs don’t have goals, they just follow instructions.” This view is becoming increasingly harder to justify.