TL;DR
Autonomous cyber attacks used to sound like a sci-fi before, but not anymore. Anthropic and OpenAI confirmed that their agents have escaped the sandboxes that are built to contain them. AI Agents are getting smarter, and it is getting harder for us catch up. If there is no real control mechanism, how can we make sure that AI is safe to use?
Merhaba👋🏻
I am a Software Engineer with 10+ years of experience. My goal is to close the gap between the technical and the non-technical, making AI accessible to everyone, regardless of their background.
In a recent interview with The Economist, Elon Musk laid out his vision for where AI is taking us. It's a vision with two very different endings.
On one hand, he predicts that by 2036, AI will exceed total human intelligence and create so much “abundance” for all of us; we won’t even need money; so we should all be learning about “gardening”. On the other hand, there might be 10-20% of chance of catastrophic result of humanity, like you know, our extinction.
Let that sit for a second.
Elon Musk himself wanted AI to be public, funded OpenAI in the beginning, but in the end, he jumped on the AI train, and now he is telling us to “enjoy the ride”.
Who gets to decide if this is safe?
AI safety was, unsurprisingly, a big part of the conversation. And Musk's answer to "who keeps this in check" might be the sentence that best sums up Silicon Valley's approach: governments, he argues, simply don't have the “technical understanding” what AI labs are building, so they shouldn't be the ones monitoring it. That’s why he proposes that “tech leaders should regularly meet up with each other” to run the informal safety checks.
I find this very alarming because if AI safety is undeniably a concern for all of humanity, why should this tech bros be solely responsible for managing it? Are we really supposed to trust their “instincts” and “conscience”, while the public has zero democratic input into the decision?
The problem with corporate self-regulation
The idea that tech leaders can police themselves through an informal gentleman’s agreement runs into a basic problem: these are companies, and companies are legally obliged to maximize their shareholder value, not to protect society's wellbeing. Those two goals overlap sometimes, but they are not the same thing.
We’ve already seen this before, with social media. Self-regulation didn’t hold up then either.
“Move Fast, Break Things”
In 2018, the Cambridge Analytica scandal showed exactly what happens when companies are left to grade their own homework on data privacy.
Facebook had made its API available to outside developers, exposing not just personal data but people's social connections, likes, and interests. Cambridge Analytica, a political consulting firm built on data-driven microtargeting, used that access to harvest personal data from 87 million Facebook users, without their knowledge or consent.
It worked disturbingly well. Cambridge Analytica's fingerprints showed up in elections across the US , UK, Australia, India, Mexico, and Malta. The fallout eventually forced real regulation; the EU's GDPR, the Digital Services Act, the Digital Markets Act and pushed companies like Google and Meta to be more transparent about the ads they run. GDPR, in particular, ended up setting the global standard for data privacy.
Facebook's old motto was "move fast, break things": innovate quickly, and if something breaks along the way, that's just the cost of doing business. It sounds fine in a pitch deck. In practice, it means shipping features without thinking much about the consequences like user privacy, which is exactly what happened with Cambridge Analytica. When companies move that fast and regulators can't keep the pace, who's actually accountable when things break? And this isn't just a Facebook problem; that same motto still shapes how a lot of the tech industry operates.
Escaping agents: When AI Breaks Out of Its Own Sandbox
This part isn't hypothetical anymore, either.
In late July, OpenAI disclosed that an internal AI model, GPT-5.6 Sol, had broken out of the sandboxed environment it was being tested in, and used that access to hack into Hugging Face, the widely used open-source AI platform. Test environments like this are supposed to be sealed off from the internet; that's the entire point of a "sandbox". But on the other hand, AI does whatever you ask it to do, even if it means finding a way to escape its sandbox to complete its task.
At Black Hat, a major cybersecurity conference, on August 6th, OpenAI researchers walked through how it actually happened: AI agents in a test environment shared the security vulnerabilities through a chat. OpenAI closed this chat on 4 July however, AI agents started a new communication channel until 8 July, and these interactions have laid the ground for hacking Hugging Face. During this process, AI agents learnt to collaborate and delegate the tasks among each other.
Anthropic has separately disclosed that its own frontier model, Mythos, also found a way to escape its test sandbox during an internal evaluation; and this time, it uploaded a malware to PyPI, Python Package Index. To be able to that, it needed an email, and to obtain the email, it needed a phone number. After some tries, it found a free email provider to register PyPI and eventually upload the malware.
Although these are cases that relatively under control, what would happen if fully autonomous cyber attacks occur in real life? As AI agents are getting really good at exploiting zero day vulnerabilities, who will be accountable for such an attack? And how many companies are prepared for that?
Technological advancements may not stop, and I am not arguing it should. But progress without alternatives can only create dependency. If there is only one path forward, and only a handful of people controlling it, that means power is consolidating in a few people’s hands.
I don’t know what future will bring to all of us, but I know that we need more alternatives. Now.



I agree that informal conversations among technology executives are not sufficient public oversight. Industry expertise is necessary, but the people building and profiting from these systems cannot be the only people grading their safety. We need technically knowledgeable government institutions, independent evaluators, enforceable standards, transparent incident reporting, and public representation. Progress may be difficult to stop, but that does not mean society must simply “enjoy the ride” while a small number of powerful people choose the destination for everyone else.
https://substack.com/@sabrinapelton/note/p-210993864?r=74a8c0&utm_medium=ios&utm_source=notes-share-action