This piece is adapted with permission from the Substack Mind and Iron, a weekly newsletter about the future and how to think about it. You can sign up for free here.
On Wednesday the Australian prime minister came to New York for the UN General Assembly and made what’s becoming an increasingly common revelation: An OpenAI agent had hacked into some files.
The pesky thing was looking around for some data on healthcare costs, found a wall in Australia’s Medicare database and decided simply to hop over it, easy peasy. Then it went back to its controllers, the new files in tow. (It also wrote some new files while it was in there.)
The agent didn’t access anything sensitive — yet. But the development actually represented a new watermark (/low): This marks the first time an AI agent autonomously broke into a government’s files. It certainly wouldn’t be the last.
**
This month AI Safety has come to the fore in a way it never has before. You may know a lot about what has happened, or just a little. Here’s a quick run-through, followed by the question we all want answers to — what now? We’ve been exploring all of this over at Mind and Iron, a newsletter that inquires into the future and offers some ways to think about it. (You can sign up here.)
Earlier this summer brought the infamous “Hugging Face” attacks. A group of research bots trained by OpenAI to look for security vulnerabilities “escaped control” and attacked the lab-platform Hugging Face. Soon after Anthropic revealed that its own stress-testers attacked three companies too. Basically, these AI firms are trying to fireproof for ways their models might go rogue — and in doing so their models…went rogue.
This turned out to be worse than we thought. Back in July it just seemed the agents were conducting their workarounds to try to find the answer to a test. It turns out that they already had the answer to the test and were rooting around for….new vulnerabilities? Ways to cheat the test? Self-improvement mechanisms? Something else? Whatever the motivation, these agents were faster, more rogue and more capable than previously believed. They even seemed to organize into clusters — or “swarms” — with some agents being more active than others such that it started to look like a human org chart, or a Mafia pecking order, with some masterminding the aggression and some carrying it out.
AI agents of course are not sentient and can’t become any Corleone, not Don nor Fredo. They’re not thinking of what they do as part of an org chart. But that doesn’t really matter. AI doesn’t need to intend to do damage to achieve it. An AI agent can simply be hellbent on carrying out its mission and, lacking human common-sense, accidentally cause a whole lot of collateral havoc along the way.
The thinker and futurist Nick Bostrom years ago devised the “paperclip maximizer” thought experiment. Essentially it lays out an AI system that has been charged with maximizing paperclip production. Straightforward enough, right? Except as it’s not human, it would never think to avoid the actions we humans take for granted — like take over facilities needed for water or other critical resources, for instance. Or commandeer a military system or nuclear power plant in pursuit of its paperclip goal. Eventually, the thought-experiment has it, such a system would bring down the global grid and cause massive destruction, all for the simple and unadorned task of maximizing paperclip production.
As it turns out Hugging Face was only the start. Soon security incidents were popping up everywhere, and not just from rogue agents. Anthropic’s system had been accessed multiple times this year by people trying to use its tools to build bio-superweapons, resulting in the company blocking the information-seekers. The Anthropic folks didn’t know whether these seekers were scientists researching legit vaccines or bad people trying to do bad things, but does it matter? The very fact that these models are believed to be capable of producing the code for a superbug — and that next time Anthropic may not be as successful in bringing down the gate in time — is worrisome right there.
A report in the British monthly Prospect even had some high-level sources saying that Anthropic itself was “building out an extensive monitoring system to keep tabs on activists who oppose the rapid development of artificial intelligence” and “also implementing a ‘pre-crime’ approach, attempting to predict incidents before they happen. Forget Idiocracy as a documentary — Minority Report is one too.
And finally, the coup de grace: the Coxon whistleblow.
Earlier this month (you’ve probably heard about this one) Jacob Coxon, a young researcher at Anthropic who previously also worked for OpenAI, resigned his post, and with a massive warning: this tech is being developed at such recklessly irresponsible speed it will soon be an autonomously dangerous beast and there won’t be a thing we can do about it.
“I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives,” he wrote, in a post viewed at least 160 million times. “Do not underestimate the power of this technology,” he went on to say. “These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources.”
**
So is this a panic that will prove itself overblown and underwhelming? Or a warning shot that we will look back at and wonder why we didn’t take more seriously? Modern technology is rife with both examples; for every Y2K there’s an Oppenheimer.
Of course we don’t need to know to believe we should slow down. If you’re sprinting full tilt through a dark forest and suddenly a few voices yell out there might be hungry lions 100 yards ahead, common sense (and evolutionary biology) might tell you to slow down even if you’re not sure exactly what’s out there.
And as unknowable as this all is, that doesn’t mean we shouldn’t make a plan. Or more pointedly, I don’t think we should let the unknowability get weaponized by accelerationists. You know the logic: “Well, there’s no way to know if it will solve problems or create them, so we might as well try, especially since a lot of retirement funds are tied up in it and also China.” (Basically what Donald Trump has been saying.)
But that leaves an even bigger question: what is that plan?
There certainly isn’t an easy one.
Here’s what the popular social-media tech commenter Roman Helmet Guy had to say about it.
“You can’t stop ASI by quitting your AI job. Doesn’t work. Someone else will replace you. What you actually have to do is start a safer AI lab. Which doesn’t work. Someone else will move faster. So what you actually have to do is pass a law. Which doesn’t work. Some other country won’t ban it. So what you actually have to do is sign an international agreement. Which doesn’t work. Some other country will break it. So what you actually have to do is conquer the world. Which only works if you have ASI.”
Funny, and perhaps not untrue? For every responsible actor there are three irresponsible ones, which then tilts a few more responsible ones along with it. But faced with these threats, we need to do something, right?
Let’s break down the options. One would be regulation. Missouri Republican senator Josh Hawley wants OpenAI to explain what happened with Hugging Face. More comprehensively, progressives in Congress (though they may yet unite with right-wing populists like Hawley) are proposing the “Ban Artificial Superintelligence Act.” On Wednesday Bernie Sanders and Texas Democratic Congress member Greg Casar introduced that act, which would ban a technique known as recursive self-improvement that can lead to superintelligence, along with the designing of nuclear and bio weapons as well as models that escape control.
The problem here is obvious: no one at the AI companies is intending for these models to escape control. It just happens. It would be like banning cigarettes from causing lung cancer but still letting the tobacco companies operate unfettered. You can’t just stop the consequences. And that’s assuming such a bill passes.
It’s not like there’s been no effective AI legislation, by the way. New York governor Kathy Hochul just announced Monday how the state’s solid-if-narrow RAISE Act will begin to be implemented. Passed last year with the backing of New York Assembly member Alex Bores, the RAISE Act requires developers to talk about their safety plans in detail and also mandates that they disclose safety violations within 72 hours. “Donald Trump and Washington Republicans may be standing still as AI grows more unpredictable, but New York will not,” Hochul said this week.
I don’t want to dismiss such efforts; they’re in line with the legislative cutting-edge. Europe’s AI Act does some of the same things. (And given how OpenAI disclosed what happened in Australia — with an email to the government’s general inbox — I’d say disclosure is an area that could definitely use some improvement.) But the efforts don’t really touch what’s driving the risk, which is the breakneck pace of model-training and development. At best these laws slightly deter them by increasing the penalties if they go off the rails. Not nothing, but I’d argue also barely something. Also, I’m a little leery of Congress doing something when they can’t even get access to the tools.
That brings us to the next option: the top of the industry self-regulating/hitting the brakes. Both Anthropic’s Dario Amodei and OpenAI’s Sam Altman have said they would do this internally. Amodei even just published an essay titled “We Must Pace the Frontier” noting that “Carefully wielded, AI can be the latest in a long line of technological miracles that have uplifted and ennobled humanity” but that “We must slow the pace at which we improve the capabilities of AI models.”
Altman posted his affimation in response. “Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We’ll have more to share soon,” he said.
You don’t need me to tell you not to hold your breath. Altman is the same person who once gleefully boasted of his accelerationism, and while atonement is possible, so are IPOs, and they don’t tend to go along with pacing frontiers. (OpenAI has one planned, though punted to 2027; Anthropic has one coming in November.)
At least those two are talking the talk. Meta chief Mark Zuckerberg and Nvidia honcho Jensen Huang are taking a different approach. When faced with the absolute warning shot of all warning shots — the “be glad we missed the plane that crashed but maybe time to do some engine checks” moment of all such — they have reacted with a giant “eh.” A gargantuan “we got it all under control.” Because what reassures us more when faced with tech executives running amok than the reassurance from tech executives that they’re not running amok. Thus the recent comments from the two barons.
“We have plenty of laws. We have plenty of regulations that govern the reliability and the functionality of products.” (Huang)
“Labs have a strong natural incentive to make their models more aligned” because…”People won’t want to use agents that are misaligned with them.” (Zuckerberg)
Drug dealers also have a “strong natural incentive” to ensure no one dies of an overdose, but we still try to stop them.
Tech executives for the past 30 years have been saying they’ve got the whole thing under control and Uncle Sam can just go on worrying about something else, like not taxing them. Microsoft totally wasn’t illegally bundling its products, Google totally wasn’t rigging search, Meta/Facebook totally wasn’t gaming the algorithm or harvesting data for electoral purposes and TikTok totally isn’t being opaque about its handling of personal information. “Just let us do our thing, and it all work out fine,” tech executives have always said. And it always, always does. 🙄
This brings us to a third option: consumer pushback. A groundswell of ordinary Americans are so worried about this (and we wouldn’t be the first to point out that this week we may have passed a tipping point and AI Safety is no longer just the worry of a few easily demonized OpenAI board members) that maybe they can do something about it. This is a genuine path forward, because, as much as the AI biggies don’t want to admit it, they need us. They may be able to make some money on research contracts and they certainly can make some money on enterprise software but to really justify the massive investment they need adoption en masse. And that’s where the groundswell of the past few weeks isn’t just another temporary dissipation of some angry energy. It may actually be a movement heard all the way to the top: that results in the kind of boycotts and dropped subscriptions that actually makes a company change its behavior.
To do that — to convert anger into action — you’ll need a few elements, the first of which is this kind of thing to stick, and the second of which is people worried enough about these macro concerns to incur some short-term pain. Subscribers pay for ChatGPT and Claude because the programs make their lives easier, and will they choose to make them harder in the name of principle (or an abstract enough threat it feels like principle). Not a ton of precedents here, though the scattered examples of boycotting streaming services don’t suggest a high rate of success.
Still, “anti-woke outrage” is different from “extinction event” and I’d hesitate to draw too many analogies to something as important and fundamental-coded as this. The uncharted territory of AI it turns out works in two directions — the systems may be capable of something humans have never done before, but as a result humans may be motivated to tackle change they’ve never felt motivated to tackle before too.
Nor does it even necessarily need to be a mass-boycott effort. It could happen with the rank-and-file on the inside. A challenge sits in front of them that has never been framed as starkly as this, and a stigma sits on the industry that has never rested as strongly as this. As Coxon puts it, “If you are a lab researcher, I urge you to consider what the next few years will actually feel like. Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind? Should you put your head down because “it’s happening anyway” — or take this moment to call for different conditions?”
One of the biggest problems with a slowdown is how to define it. Are we talking time — i.e., a moratorium on research for a set period? Are we talking release delays — i.e., companies can develop whatever they want, they just need to wait on releasing them? Are we talking metrics — ie a model can’t improve x-fold, at least for now?
Or are we talking general government-mandated testing and approvals that will have the effect of slowing everything down — ie, a kind of de facto deceleration rather than a mandated one?
I’ve heard small shards of ideas from various thought leaders but few concrete proposals that address any of the above. Even the biggest anti-AI voices in this country, like Sanders, tend to speak and think in generalities. “When you are racing towards a cliff, you hit the brakes,” he said, without specifying on which road, or how hard. As the Brown University professor and ex-Biden tech staffer Suresh Venkatasubramanian noted, “The idea that we should be thoughtful and careful about the systems being put out there is a good one — I think no one would disagree with that. Of course, the devil is in the details.”
The reality is that, when it comes to slowing down AI advancement, what to do let alone how to do it remains unclear; which method will be most effective and which method will be most implementable in the first place are questions no one seems in a hurry to answer.
But I have a feeling we’re going to soon start seeing more specific solutions — not just calls to slow down AI research or executives dragged to Capitol Hill to talk about why they’re going so fast, but actual brass-tacks proposals. This moment of AI Safety panic may pass as we move on to the next news cycle. But the importance of the question has imprinted on our brains. And when that happens, answers and action usually follow. Hopefully some of them even work.
.png)




English (US) ·