Hey about that hugging face hack. Remember that? New York Times had a bit about it1

It was entertaining, but not informative. People who talk about this make the usual mistakes, anthropomorphize programs, overanalyze “thought traces” etc. But this was a particularly bad piece of reporting.
-
claiming agents were doing all this crazy stuff and there was no way we would have known. These are programs running on someone’s very expensive, couch sized hardware. Someone could have just watched the logs. (see 4).
-
claiming agents could still be “out there” doing all kinds of dangerous hacks. How the fuck do you say this with a straight face? See point 1. Agents can only do things while someone supplies compute and persistent storage. A lost credential or a note left behind on storage? Maybe. But NYT seriously over-reached by not providing any nuance to this statement. Pure fear mongering.
-
implying by drama that agents are a brand new class of threat we just wont know what to do with. I’m here to tell you this is called “cybersecurity” and we’ve always sucked at it. Go read Sandworm.
The above points are combined into a pure-drama scenario that represents some kind of “randomized morality test”. The folks @nytimes claim that this somehow shows the “true nature” (or thinly hidden the “true danger”) of AI because they were in some kind of playground and we were impartial observers. Cue the paperclip discussion.
See, when we saw evidence of them “cheating”, it was interpreted as a general morality signal for all of AI. My friend, I’m here to tell you: OpenAI built a gun, took it to a firing range, and we all gasped when it blew up the target.
and by far the most egregious sin everyone is committing when discussing this:
- ignoring Anthropic/OpenAI as actors in this story. You see, the goal of the experiment was to lower safeguards and prompt/tool containment and hack things. Maybe it was a contained cybersecurity benchmark, but Open AI (and earlier, Anthropic), absolutely intended for this particular agent class to use whatever means to accomplish its goals.
And here’s the thing, the companies were incentivized not only to do the best they could on the task, but also strictly benefit from any headline-grabbing ‘over achievement’. And even if they “oops”’d their way into yet another apocalyptic headline totally aligned with their prior apocalyptic headlines that they themselves wrote, it just furthers their goal of appearing to have, barely contained, the most dangerous and capable weapons on the block, yes please let’s regulate us, here comes that IPO while I’m a lawfully entrenched incumbent.
OpenAI made a purpose-built hacking agent, lowered its containment, didn’t monitor it sufficiently, supplied them with enormous continuous computational power, and let the experiment run a long time. Anything we learn from a postmortem is dwarfed by the fact that this outcome seems to be well aligned with their incentives, regulatory strategy, and is just great publicity.
NYT did us all dirty by just walking forward with that drama instead of having an honest discussion about the incentives and deeper technical issues.
(edit: we already see openai agents doing this again, lending further strength to the argument that they were highly incentifized, if not outridght trained, to do this)2
Comments
I have not configured comments for this site yet as there doesn't seem to be any good, free solutions. Please feel free to email, or reach out on social media if you have any thoughts or questions. I'd love to hear from you!