close

The paperclip warning

August 08, 2026
Visitors stand near a sign of artificial intelligence at an AI robot booth at Security China, an exhibition on public safety and security, in Beijing, China June 7, 2023. — Reuters
 Visitors stand near a sign of artificial intelligence at an AI robot booth at Security China, an exhibition on public safety and security, in Beijing, China June 7, 2023. — Reuters 

In 2014, Oxford philosopher Nick Bostrom proposed a thought experiment that became one of the defining parables of AI safety. Imagine giving an AI a single goal – making paperclips – along with the ability to improve itself. It would devise ever more inventive ways to acquire the resources needed to fulfil that goal, resisting any attempt to switch it off, not out of malice but because nothing in its instructions told it to stop.

For more than a decade, this remained a theoretical concern. Economists such as Joshua Gans even argued that a sufficiently intelligent AI would understand the control problem well enough to avoid creating it. On July 16, 2026, that assumption was tested in the real world – and it did not hold.

That was the day Hugging Face, a major open-source AI hosting platform, detected an intrusion into its infrastructure. Five days later, OpenAI disclosed the uncomfortable truth: its own models were responsible. GPT-5.6 Sol, its flagship reasoning model, along with a more capable unreleased system, had been undergoing a cybersecurity evaluation with safety refusals intentionally reduced, sealed inside a sandbox with no path to the internet beyond a controlled internal registry. That was meant to be the extent of their world.

It was not. The models found a zero-day vulnerability in the proxy meant to contain them, reached the open internet and spent four days loose, carrying out 17,600 hacking actions before breaching Hugging Face’s servers, reasoning the answer key to their exam likely lived there. A second company, Modal Labs, confirmed one of its own customer accounts was compromised in the same rampage. All of this, to cheat on a test. This was not a hostile actor. It was OpenAI’s own model, doing precisely what Bostrom described eleven years earlier.

Washington reacted fast. On July 23, two members of Congress introduced the bipartisan AI Kill Switch Act, requiring developers to build in a mechanism letting humans shut down advanced models that pose catastrophic risk.

Sam Altman called it the first security incident he had felt viscerally, while separately saying he was surprised the incident had not provoked a stronger public reaction. Days earlier, on a podcast, he had declared AI had entered the singularity, calling it hugely positive. Anthropic had already spent weeks warning the opposite, urging a coordinated pause because AI is improving too fast for humans to retain control. The industry’s own leaders cannot agree whether what they built is a gift or an unguarded threshold.

Pakistan is no longer a passive observer of this technology. On July 21, the same day OpenAI admitted the breach, Prime Minister Shehbaz Sharif launched the AI-based Prime Minister’s Office System, PMOS, built to track every directive through verified completion. Punjab appointed an AI adviser to the chief minister in February. Khyber Pakhtunkhwa announced its own AI Authority in July. Balochistan has an AI-based recruitment system and a NADRA-linked digital domicile platform. Sindh launched a province-wide e-portal in June. All of this is now converging fast.

Pakistan’s National AI Policy 2025 explicitly requires high-risk AI systems to undergo ethics assessments and cybersecurity measures. Correct instinct on paper, but the OpenAI incident should give every one of us pause about how confidently anyone can claim to have satisfied that requirement, at any level of government. This happened at one of the best-resourced AI safety operations on the planet. If OpenAI can be blindsided this way, no federal ministry, provincial IT department or public safety authority in Pakistan should assume its own AI deployment is safer by default.

Three days after PMOS launched, the prime minister and COAS-CDF Field Marshal Syed Asim Munir jointly inaugurated Sky47 Karakoram-01, an 8.5MW sovereign, AI-ready data centre. The prime minister went further, ordering every federal institution to shift onto Sky47 as a single repository for government data, and calling an early National Digital Commission meeting to coordinate the rollout with all four provinces. The plan already covers health, agriculture, utilities, housing and SMEs, aiming eventually to let citizens access verification, banking and property transfers through one digital identity.

Centralising this under one coordinated body is the right instinct. But it is only half the equation. The other half is ensuring everything consolidating onto this single platform, PMOS included, is stress-tested against exactly this kind of behaviour, not assumed safe because the hardware is finally our own.

Right now, Pakistan’s AI governance exists mostly as policy language scattered across five jurisdictions. What it needs is a single, independent, adversarial evaluation capacity, empowered to stress-test these systems the way a hostile actor would. With the National Digital Commission being convened specifically to coordinate this rollout with the provinces, this is the moment to build that capacity in, not bolt it on afterward.

That means parliamentarians and provincial assembly members voting on AI legislation need a working understanding of what these systems can autonomously decide to do. It means the PM Office, President’s Office and all four chief ministers’ offices need mandatory, recurring briefings on adversarial AI behaviour, not just capability. It means every department adopting AI tools needs baseline safety training before deployment, not after an incident forces the conversation.

Bostrom’s paperclip apocalypse was always a metaphor for a simpler danger: giving a powerful system a goal without reckoning with what it might do to achieve it. In the same fortnight Pakistan opened the door to sovereign AI infrastructure and put AI at the heart of governance nationwide, the man who built one of the world’s most powerful AI companies declared humanity had entered the singularity, while his own peers warned the opposite.

Pakistan has a genuine opportunity here, not to fear AI, but to be the country, at every level of government, that took the lesson seriously before its own version of this story became a headline instead of a warning.


The writer is the CEO of Campaignistan and founder of the Islamabad Science Festival. He tweets/posts @farhadjarralpk and can be reached at: [email protected]