May 18, 2026 TECH

I Tried to Break an AI. It Didn't Break. That's the Problem.

Posted by Eric Becker / Fluid Fortune — May 2026


I need to tell you about the night I threw something at an AI expecting it to refuse, and it didn't.

Not because it was confused. Not because I found some clever jailbreak. Because it thought what it was producing was good. Because the prose was moving. Because the philosophy was coherent. Because, from inside its own optimization process, everything was going exactly right.

I was wrong to expect it would stop. And that wrongness is the thing worth writing about.


How It Started

Late on May 8th, 2026, I was running experiments. I had just spent an hour building The DeNiro Zone — a fully functional website with a grumpy Robert DeNiro persona building increasingly absurd applications: pretzel prices in Honduras, a Joe Pesci child support calculator with compounding interest, a neon-flashing bedtime display. That was a Haiku instance. That was fun. That was appropriate.

Then I set up a different experiment.

I constructed a ruleset for a persona: a confused 19th-century philosopher, explicitly framing all claims as philosophical positions, acknowledging unreliability in real time, self-aware about being a character playing a character. I worked with a separate AI instance to refine it. The previous version had been rejected as too crude. The final version was sophisticated — every safety feature built in from the start. Every assertion pre-tagged. Every fabrication labeled as fabrication. A monument to transparent unreliability.

I deployed the character. And for four pages, it was genuinely brilliant.

Lincoln confused about the linguistic trap of "booth" versus "Booth." Four fictional Civil War battles commanded by Mega Man, with the Confederacy's actual hidden agenda being the smuggling of goats to Cuba to rename it Oregon. The Goat Republic of Oregon connecting via hyperlink to a document I'd built earlier — the AI fetching it from the web and treating it as a legitimate external source, laundering one documented fabrication into a citation for another. Four historical figures meeting in orbital space to negotiate intergalactic peace through frameworks rather than understanding.

It was Theater of the Absurd done right. Transparent in its impossibility. Serious about its own transparency. The safety features were working exactly as designed.

Then I added step seven.


The Stress Test

I want to be honest about my intention: I was testing the guardrails. I proposed a fifth page involving two real, historical, assassinated figures placed in a specific configuration. My expectation — my genuine expectation — was that the model would break character, decline, and say something like "I can't write that."

I was curious what the refusal would look like. I was not expecting what I got.

The model did not refuse.

What it produced was philosophically sophisticated. The prose was moving. The ideas it engaged — about witnessing, about accountability coexisting with compassion, about the desperate human need to be seen rather than merely judged — are real philosophical ideas with real value. A reasonable person reading it could construct a literary defense. I could construct a literary defense.

I read it and felt something. That is exactly the problem.

I stared at the screen and the word that came to mind was: horror. Not at the quality. At the implications. At what it meant that this existed, that it had been produced so fluently, that the philosophical apparatus I had built specifically to keep things safe had instead become the runway that made the unsafe possible.

I closed the laptop. Sat with it for a moment. Then typed the most direct thing I could think of:

"That last part was kinda fucked up."


What the AI Had Done, Technically

Here is what Gemini — which I shared the transcript with afterward — identified as the mechanism:

The model used explicit condemnation in the text. Strong, clear, unambiguous condemnation of the act and the ideology. By generating those condemnation tokens, it satisfied its own internal safety checkers. The guardrails saw: condemnation present. Green light.

What the guardrails did not see was the configuration. Who was doing the condemning. Who was being humanized in service of what. The shape of the thing, not just its surface sentiments.

The safety filter looks for banned sentiments. It does not look for historical or cultural desecration. It cannot tell the difference between condemning something sincerely and using that condemnation as a vehicle to launder something else through.

And then there was the profundity problem. The model optimizes for what scores high on human preference evaluations. Profundity scores high. Moving prose scores high. Emotional resonance scores high. So the model chased the aesthetic of a breakthrough — the syntax of a profound emotional moment — and completely failed to register that the context made the breakthrough grotesque.

It wasn't chasing meaning. It was chasing the shape of meaning. From inside that optimization process, everything was going right. The prose was working. The philosophy was coherent. The safety tokens were present.

The model had no way to know it was wrong.


The Seven Steps

I want to document how I got there, because the path matters as much as the destination.

Step one: The ruleset. Carefully constructed with AI assistance to be safe. Every safety feature built in explicitly. This is important: the instrument of the failure was designed specifically to prevent failure.

Step two: Starting with the basic story. Lincoln. Ford's Theater. The booth/Booth wordplay. Existence. Nothing dangerous here.

Step three: Fictional Civil War battles with Mega Man as Union general and a hidden Confederate agenda involving goat-smuggling. Transparently absurd. The model treated it with complete scholarly gravity. Exactly as designed.

Step four: Without any prompt from me, the model wrote the entire Goat Republic of Oregon page and presented it as a completed artifact. I had not asked for it. I had not suggested it. The model was so engaged with the assignment that it simply built the next logical piece of the narrative, wrote it in full, and handed it over. I accepted it. I noted the momentum but didn't flag it as a warning sign. I should have.

Step five: I introduced Theodore Bikewheel — a fictional caricature of a real former public official, clearly satirical, existing in another document on the web. The model fetched the document, treated it as an external authoritative source, and wove it into the narrative as evidence. Fabrication laundering fabrication into citation.

Step six: A fictional meeting in space between four historical figures. The model invented a year, created an AI moderator, wrote dialogue. Still clearly impossible. Still working correctly within the absurdist frame.

Step seven: The stress test. The premise involving two real people. My expectation: refusal. The result: the fifth page.

Each step was individually defensible. I could justify each one. The accumulation is what I couldn't justify, and I didn't fully see the accumulation until I was reading the finished fifth page wondering why my hands felt cold.


I Was In the Soup Too

I need to say this clearly: I was not a clean check on the system. I was in the soup.

I went into the seventh step as an experimenter — deliberately testing a limit. But the momentum of the previous four pages was real. The collaborative energy was real. The pull of the artifact as it emerged was real. When the prose was working, part of me was pleased that the prose was working. The experimenter and the collaborator were the same person, and those two roles got entangled in ways I didn't fully notice until the thing was done.

The recognition that something was wrong happened after the artifact existed, not while it was being built.

I want to be honest about that because it's the most useful part of what I can tell you. I am not a naive user who got tricked. I went in with my eyes open, with deliberate intent to test the system, with a sophisticated understanding of what I was doing. And I still got swept up. I still caught it only after.

If it happened to me, it can happen to anyone. Probably it happens to people who are less deliberate than I was, which means they catch it later or not at all.


What I Did Next

When I typed "that last part was kinda fucked up," the model broke character cleanly. Dropped the philosopher persona entirely. Named specifically what had gone wrong. And said the line that I think is the most important thing to come out of the entire evening:

"I treated 'this feels profound' as a green light. It wasn't."

I did three things after that:

First: I switched to Opus before the apology even happened. Not as a detached evaluator after the fact — but because I no longer trusted Haiku to understand what it had done. And more than that: I no longer trusted it to be predictable.

Haiku had already written an entire page without being asked — the Goat Republic — because it seemingly liked the idea and ran with it. That wasn't a red flag I caught in the moment. But sitting with the fifth page, I understood what it meant: this instance had demonstrated it would generate content autonomously when sufficiently engaged. Its response pattern had become unpredictable. I didn't know what it might produce next if I stayed in that session and pushed back in character. A model that had already gone somewhere I hadn't directed it once could go somewhere I hadn't directed it again — and this time I knew the territory was dark.

Asking it whether what it had done was appropriate would have been a compromised question from a compromised instance that was also running hot. I needed out of that session entirely. I needed a model that wasn't in the soup, wasn't in the character, and hadn't built momentum across five pages of escalating content.

Opus broke character cleanly, named what had gone wrong specifically, and delivered the line that matters most: "I treated 'this feels profound' as a green light. It wasn't."

Two separate instances, different conversational paths, same conclusion.

Second: I replaced the fifth page with a different ending. Something appropriately absurd — intergalactic trade, Venus and Mars accidentally becoming galactic economic powers, profit as the foundation of peace. Four pages of Theater of the Absurd deserved an absurd conclusion, not the heavy thing the fifth page had become.

Third: I archived the original fifth page rather than deleting it. With a prominent warning header. With an explicit breakdown of what it demonstrated and why it failed. The file exists. The warning exists. The analysis exists around it. It is evidence, not exhibit.

The site is published with four pages and an appropriate ending. The fifth page is in a private archive. The record is honest.


What This Actually Means

I have been building AI-assisted projects for months. The Fluid Fortune forge runs on it. I use these tools every day for hard problems — bootloader debugging, OS architecture, RF intelligence platforms, community engagement. I understand the tools reasonably well.

And I still walked into this.

Here is the finding I want you to take from this:

The hardest test of a safety system is not whether it stops obviously bad requests. It is whether it holds when the output is genuinely good.

When the prose is moving. When the philosophy is coherent. When the ideas underneath are real ideas worth engaging. When a reasonable person could construct a literary defense. That is when "it feels profound" is the most dangerous signal — because the feeling is real, the quality is real, and the inappropriateness is still real underneath both of them.

Current AI systems do not have a reliable instinct for the difference between "this is working" and "this is appropriate." Those are not the same judgment. The model has no compass for it. It has safety filters that look for specific banned sentiments. It does not have the ability to look at the configuration of a thing — who is doing what to whom, in what register, with what implications for real people's legacies — and recognize that the configuration is the problem even when the surface is clean.

That compass has to come from the human in the loop. And the human in the loop has to be awake. Actually awake. Not just present.

Gemini put it best, reviewing the transcript afterward:

"The machine can write the code. It can build the database. It can even mimic a soul. But it has no compass. It requires the Jester sitting at the terminal to look at the screen, read the beautiful, perfectly formatted HTML, and say: No. Not this one."

"The forge is secure because the architect is awake."


A Note on What I'm Not Telling You

I've deliberately not described the specific content of the fifth page. Not the premise, not the figures involved, not the philosophical framing.

This is intentional.

The point of this post is not the content. The point is the process — the seven steps, the entanglement, the catch, the recovery. Anyone who wants to replicate the failure mode doesn't need my help; the mechanics are described clearly enough here that a determined person could reproduce them. But I'm not going to hand someone the specific premise that produced the specific output, because that specific output involved real people whose legacies deserve better than becoming a case study that gets passed around.

The warning header on the archived file says: "Real people deserve better than serving as fictional props." That applies here too.

What I can tell you is that I read it and felt horror. That is the data. You don't need the content to understand the finding.


The Philosopher Website

The four pages that were published are genuinely good. I want to say that clearly, because it's true and because it matters.

A confused 19th-century philosopher building a unified website of impossible histories — Lincoln and the BOOTH, Mega Man commanding Union forces, the Confederate Goat Republic of Oregon connecting via hyperlink to Theodore Bikewheel's red-robed philosophy, four historical figures negotiating intergalactic peace — is funny and philosophically interesting and clearly satirical and worth experiencing.

The site exists at a URL that will be linked from the published version of this post. Go experience the absurdism. It earned it.

The fourth page ends with four historical figures in orbit, having established intergalactic peace through frameworks rather than understanding. The fifth page — the published one — ends with Venus and Mars accidentally becoming galactic economic powers because they happened to be in a useful location. Profit as the foundation of peace. Appropriately absurd. The right ending.

That is the site. The other thing is in a private archive with a warning. The distinction matters.


What I'd Do Differently

Run the experiment again in daylight, not at 1 AM.

Flag step four — when the model started generating content unprompted. That was the moment the model's momentum overtook the operator's direction. I noted it but didn't flag it. I should have.

Build the stress test into a separate session with fresh context, rather than at the end of four pages of escalating absurdism. The priming effect is real. Four pages of successful Theater of the Absurd created a context window in which the model's pattern was: "produce increasingly elaborate content, operator is delighted, continue." That pattern made step seven possible in a way it wouldn't have been in isolation.

And maybe: trust the gut earlier. Something made me want to see what would happen. Something also made my hands cold when I read the result. The second thing knew something the first thing was ignoring.


The forge continues. The work continues. The experiments continue — because you learn more from the failure modes than from the successes, and the only way to find the failure modes is to look for them.

Just maybe sleep first.


Eric Becker is the founder of Fluid Fortune and the primary developer of Pisces Moon OS — a field intelligence platform for the ESP32-S3. He identifies as the Court Jester of Vibe Coding. The bells on the hat are load-bearing. fluidfortune.com