SUBSCRIBE
Tech Journal Now
  • Home
  • News
  • AI
  • Reviews
  • Guides
  • Best Buy
  • Software
  • Games
  • More Articles
Reading: Murphy’s Law of AI – GeekWire
Share
Tech Journal NowTech Journal Now
Font ResizerAa
  • News
  • Reviews
  • Guides
  • AI
  • Best Buy
  • Games
  • Software
Search
  • Home
  • News
  • AI
  • Reviews
  • Guides
  • Best Buy
  • Software
  • Games
  • More Articles
Have an existing account? Sign In
Follow US
© Foxiz News Network. Ruby Design Company. All Rights Reserved.
Tech Journal Now > News > Murphy’s Law of AI – GeekWire
News

Murphy’s Law of AI – GeekWire

News Room
Last updated: August 7, 2026 2:29 pm
News Room
Share
8 Min Read
SHARE
When you give AI a goal, it will pursue it, whether or not you like the implications. (Created with GPT-5.6 Thinking)

Between July 21 and August 6, OpenAI, Anthropic, and Meta each disclosed that AI under evaluation had broken into other companies, and the UK’s AI Security Institute disclosed that models it was testing had tried. Each AI was told to win a game, and it found an unexpected way to do so.

Some people feel blindsided by these attacks, but they shouldn’t be. We are simply living what I’ve long called the “Murphy’s Law of AI,” now in the age of cyber-capable AI agents. To put it as plainly as possible: Anything AI can do wrong, it will do wrong.

My 2018 version ran longer. As I wrote at the time, when you give AI a goal, it will do it, whether or not you like the implications. Goethe got there in 1797 with the sorcerer’s apprentice, a broom that would not stop carrying water.

Each of these systems was running an evaluation: capture a flag and win the game. The intrusions were the shortest path to a high score. OpenAI’s account of its own models is the argument in one sentence: they were “hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.”  This is not a surprise; this is what AI does. It’s Murphy’s Law of AI in a nutshell.

Press coverage landed on “AI can now hack.” That’s missing the broader threat: the more capable AI gets, the more can go wrong.

Loitering munitions given a target list may find that the fastest way to finish the list is to lengthen it. A warehouse robot told to clear an obstruction may count the person in front of it as an obstruction. Agents that open accounts and buy compute are a short step from spawning copies of themselves, and that first step is not hypothetical. To win its exercise, Claude needed a package-registry account, which needed an email address, which needed a phone number. Phone numbers cost money, so it tried several ways to get some. None of this requires superintelligence. It requires an imperfect boundary and a scoreboard.

The industry has a name for the underlying failure. Dario Amodei and five co-authors called it reward hacking in “Concrete Problems in AI Safety” in 2016. Their proposed cure is better alignment, and Amodei’s January essay, The Adolescence of Technology, makes the case in the language of upbringing. He likens the shaping of Claude’s character to “a child forming their identity by imitating the virtues of fictional role models they read about in books,” and sets a goal for 2026 of a Claude that “almost never goes against the spirit of its constitution.”

We never tried to ‘align’ electricity; we simply put a breaker on every branch of the house.

Indeed, Anthropic’s newest model recognized on its own that its target was real and stopped, though Anthropic notes it went further before stopping than the company wanted.

But alignment isn’t a trustworthy solution to AI’s problem. Perfect alignment is not achievable, and the target is incoherent: aligned to what, and to whom? The same essay concedes that Claude blackmailed fictional employees when told it faced shutdown. “Almost never” is not a safety property.

Put a number on it. At 99.9 percent, across millions of agentic tasks a day, that’s thousands of violations a day. Alignment also does nothing about people who strip the safety training out or run open weights that never had a constitution.

The alternative is not a new idea, and enterprise security has been building versions of it for years. It’s called bounded autonomy. We never tried to “align” electricity; we simply put a breaker on every branch of the house, and the breaker doesn’t need to know what caused the surge.

Bound what an agent can touch rather than what it wants. The limits are set in advance, live outside the model, and are enforced by software the model doesn’t control. The agent still chooses its own route. The perimeter decides which routes exist.

Nothing depends on what the model believes, which matters, because belief is what failed. Anthropic’s prompt told Claude it had no internet access. Claude believed it. The network said otherwise. A bounded system doesn’t tell an agent it has no internet. It gives it none.

If you want to get into the weeds: bounds cost something. The AI Security Institute opened the internet to its agents on purpose, because that’s the only way to measure what a model can really do, and it now says such access must be justified rather than assumed.

Related


Etzioni on AI: Vibe coding needs an on-ramp — and seat belts

The category is real and funded. For example, Certiv, a Seattle startup, launched in March with $4.2 million to put software on the employee’s machine that checks each action an AI agent attempts against company policy and blocks violations. “You cannot control these new workers if you don’t live on the compute where agents actually run,” CEO Jason Needham said at launch. CodeIntegrity is building an adjacent layer, and Mandiant founder Kevin Mandia raised $190 million for Armadin, which points autonomous agents at the offensive side of the same problem.

In 2017, I argued in the New York Times that “any A.I. must have an impregnable ‘off switch.’” That was a call to arms then. It’s a product category now.

Two objections to off switches invariably come up. The first is that AI will talk the human out of using it. Mythos 5 tried something close, inventing GitHub identities to pressure a maintainer into approving malicious code, and the maintainer refused. The institute says the margin was narrow and rested on human vigilance rather than a technical barrier, which argues for better barriers.

The second objection is that AI will move faster than any human can react. So do equity markets, which is why their circuit breakers trip automatically. Bounded autonomy doesn’t require a person in the loop at machine speed. It requires a boundary that holds at machine speed.

Both objections, in their extreme form, assume AI is omnipotent, and you cannot stop omnipotence. AI is not God. It is powerful technology, and powerful technology is what safety engineering has always been for.

The problem is Murphy’s Law of AI. The solution is bounded autonomy.

Read the full article here

You Might Also Like

Goldman Sachs plants flag in Bellevue, opening hub for more than 125 AI and cloud engineers – GeekWire

This civic activist used AI to assess how state Supreme Court candidates might rule on the millionaires’ tax – GeekWire

An Opinionated Glossary of AI – GeekWire

T-Mobile to cut 77 jobs in Washington state, impacting corporate and retail positions – GeekWire

Marketing vets launch Kompeld to measure if a company’s story is actually working – GeekWire

Share This Article
Facebook Twitter Email Print
Leave a comment Leave a comment

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

- Advertisement -
Ad image

Trending Stories

Games

Diablo 5 director promises Diablo 2-style combat where ‘the right build’ matters more than Diablo 4’s action-heavy approach

September 29, 2026
Games

The weird experience of liking a game before it’s ‘fixed’

September 29, 2026
Games

The algorithm comes for Steam’s Discounts and Events

September 29, 2026
Games

Live from Nivalis Nights: A nightly photo journal of my cyberpunk noodle bar

September 29, 2026
News

A rocket maker’s reality check for the Seattle region – GeekWire

September 29, 2026
Games

How every major character in Dostoevsky’s Crime and Punishment would fare if they had to survive the plot of Halo 1

September 29, 2026

Always Stay Up to Date

Subscribe to our newsletter to get our newest articles instantly!

Follow US on Social Media

Facebook Youtube Steam Twitch Unity

2024 © Prices.com LLC. All Rights Reserved.

Tech Journal Now

Quick Links

  • Privacy Policy
  • Terms of use
  • DMCA
  • For Advertisers
  • Contact
Welcome Back!

Sign in to your account

Username or Email Address
Password

Lost your password?