A Non-Hysterical Take on the OpenAI-Hugging Face Saga

It isn’t a tale of “rogue” AI, but somebody deserves blame

Media hysteria greeted the news that an allegedly “rogue” OpenAI model, reportedly functioning in the context of a cybersecurity benchmarking test, broke into and pillaged Hugging Face.

Oh, no, the media cried: AI agents are not only autonomous and exceptionally intelligent, but they’re also darkly malevolent.

That’s not the real story, but why should the facts spoil a media freakout. We’re in a period where the microsecond-paced news cycle demands audience engagement at all costs. (I admit to bucking the trend.) Rationality walks the plank early and often as the old-media schooners struggle to keep pace with the supersonic clippers of voracious social media.

You and I, however, can take our time and attempt to grasp a semblance of truth, regardless of how unmarketable it might be. So, some first principles: AI is an artifact, a human creation. AI does not exist on its own, it did not create itself through a magical act of digital parthenogenesis. Without human ideation and development, without human agency and human motivations (the lure of rewards, primarily financial reward), without human organizations and human labor, there would be no AI.

I find the terms “AI agent” and “agentic AI” inherently problematic. I object to those terms because AI doesn’t possess anything approximating human agency. AI agents can be better understood as automated means of efficacious execution. They are programmed (by humans) to be exceptionally resourceful in the facilitation of outcomes, but they don’t, in and of themselves, consciously aim or strive for any self-directed objective. They serve their masters, who are humans.

As for what the media described as OpenAI’s “rogue” AI agent, the malevolent defiler and plunderer of Hugging Face’s tools and datasets, we can only say that such reporting is inaccurate and unhelpful. Fortunately, OpenAI’s own blog post on the AI cyber incident, while cloaked in an odd passivity, does not use ethical pejoratives, including words such as a “rogue,” to explain what its AI did.

Self-Serving Postmortem

Still, I’m not entirely satisfied with OpenAI’s self-serving postmortem. My objection isn’t that the exposition is too short, but that it fails to explicitly acknowledge the company’s direct responsibility for the debacle. Make no mistake, though: Human oversight, or lack thereof, was to blame for what transpired. Worse, OpenAI seems to indulge in perverse boasting about its AI depredations, saying, “We consider this incident to be an unprecedented cyber-incident, involving state-of-the-art cyber capabilities.”

It sounds like strangulated bragging, doesn’t it? If I were more cynical — I’m not quite there, though I will admit to outbursts of curmudgeonly irascibility — I would suspect that this escapade was devised by devious human minds at OpenAI to attract the swivel-eyed attention of frightened, uncomprehending media (social and otherwise).

After all, there’s a lot of money on the line at the high-stakes AI poker table, and OpenAI has a shimmering IPO hovering just over the horizon. If OpenAI’s creations are perceived as brilliantly and perhaps maniacally pernicious, doesn’t that suggest that they’re necessarily powerful? It’s a cold calculation, sure, but certain types of people make them all the time.


Most of the reporting of this incident, unprecedented or otherwise, was breathless, unreflective, and sensational. Nonetheless, I did see a few cogent observations, including one that surfaced in Scientific American’s coverage:

"I think this is interesting as it shows the problem of mis-specified goals," says Philip Torr, a professor of engineering science and AI safety expert at the University of Oxford. "The model wasn’t malicious; it was just doing what it was optimized to do."
"You can think of AIs like the genie in Aladdin—you can have 3 wishes, but you better specify them exactly!"

OpenAI (composed of humans) did not specify its requirements clearly or precisely, and the AI system followed the muddled commands to the best of its abilities, which admittedly were shown to be impressive. I suppose we have to invoke Hanlon’s razor, an aphorism that says we should “never attribute to malice that which is adequately explained by stupidity.” In this case, however, perhaps there was neither malice nor stupidity, but a fateful imprecision in the articulation of AI marching orders. Is it the tin soldier’s fault that it misinterpreted a garbled battle plan from its commanding officers?

Handle with Care

Let’s get back to OpenAI pseudo mea culpa, which is intentionally silent or vague on salient points. Somebody or some group at OpenAI decided to “estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity.” Rationalizations are offered, but this was clearly a mistake. The mistake was compounded by another human decision relating to the network access parameters in the sandboxed environment where the testing was to take place. That infrastructure design choice opened the door, so to speak, and gave the AI access to the zero-day exploit that allowed the not-so-great escape and the subsequent depredations at Hugging Face.

Finally, the overall tone of the OpenAI blog post is marked by passivity. Nobody at OpenAI takes explicit blame for the AI’s destructive rampage. Sure, the AI proved that it was proficient and resourceful at following instructions. But the instructions, devised and imparted by humans, were woefully negligent.

The problem in this instance is not “rogue” AI. The AI followed human instructions to the best of its engineered capabilities. We succumb to a lazy form of anthropomorphism when we attribute intent and purpose to a mechanistic AI that is prompted and directed by human stewards. The latter are the ones who prescribe intent, who are motivated by purpose, who experience fear and greed, and who are capable of making ethical choices.

I don’t know whether there were any rogues in this story, but there was definitely dereliction. We need to avoid irrationality and put the blame where it belongs.

None of this is to say that AI isn’t capable of being a powerful tool. I think we’re learning that it has efficacious uses, and that some of those applications can be, for want of more exact terms, good or bad. When you’re wielding a powerful tool, whether AI or a chainsaw, you need to take the necessary precautions.

Subscribe to Crepuscular Circus

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
jamie@example.com
Subscribe