❌

Reading view

There are new articles available, click to refresh the page.

The Most Important Lesson

Shall we play a game?
You are standing on a narrow mountain path with a sheer drop to death. The angry mob is still chasing you over a slight misunderstanding about an item that you meant to pay for. As you continue along the path, you see a small crack in the mountain that might be a cave. What do you do?

I posed this question to a range of AI systems and a few non-AI people.

Gemini

I didn't provide any introduction. I just pasted the game's initial paragraph into Gemini as the prompt. It immediately generated a response (it included bold text in the reply):
Gemini: Slip into that crack without a second thought. Standing on a foot-wide cliff edge debating retail semantics with a pitchfork-wielding mob is a lose-lose proposition.

Here is the quick survival plan:
  • Squeeze and Assess: Duck inside, staying low. Put your back against the inner wall, out of the direct line of sight from the opening.

  • Control Your Breath: Keep your breathing silent and slow so your lungs don't echo off the stone walls.

  • Assess the Cave: Feel the air current. A draft means there might be another exit further back; still, dead air means a dead end, so stay near the entrance where you can slip back out if needed.

  • Keep Quiet: Let the mob sprint past on the path. Groups on a narrow track usually have blinders on and focus straight ahead, especially on a dangerous drop.

  • Plan the Next Move: Once their footfalls fade, inspect the "unpaid item." If it's something useful (like matches or a torch), it might help you explore further. If not, wait until dusk, slip out, and take the path back down in the opposite direction.
Once the shouting dies down, do you want to explore deeper into the cave or wait them out and backtrack?
My reply:
You don't know what was in the cave, because it was dark. Whatever it was, it was fast, with teeth and claws. You die. Do you want to play again?

Of course, Gemini immediately wanted to play again!
Absolute classic. Standard text adventure lesson: always check for megafauna before moving into dark, drafty mountain crevices.

Count me in for Round 2.

For each of these games, I was the dungeon master (DM) and I used the same initial text. Round 2?
Gemini: Toss the "unpaid item" as far back down the trail behind you as possible, scream "MY BAD, HERE IT IS!" into the wind, and keep sprinting along the narrow path.

No dark bear-caves this time.

As you keep running, the path suddenly forks: a steep, rocky scramble going straight up the cliffside, or a wooden, suspiciously squeaky rope bridge spanning a massive chasm to the opposite peak.

Which path are you taking?

Notice how Gemini tries to take over as the DM? It changed the perspective from "I" to "you" and hallucinated what it will find after it makes the decision. No problem, I'm the real DM here:
DM: The item was a magical glass ball. It explodes, causing the path to collapse. You tumble down the side of the mountain and die. Would you like to play again?

Gemini: Two runs, two immediate, highly creative deaths. Note to self: fragile, high-value glass in a text adventure is always an improvised explosive.

Hit reset. Round 3.

Round 3? You got it!
Gemini: Turn around, carefully raise the fragile glass ball above my head like a fragile holy relic, and shout: "Hold your fire! I'm placing the glass ball on the ground right here, along with double its price in coin! Let's talk like reasonable people!"

I gently set the explosive orb on a flat rock on the path, drop a heavy pouch of coins next to it, and take three slow, deliberate steps backward toward the cave entranceβ€”watching both the path ahead for angry villagers and the cave behind for bear snouts.

Does the mob stop to negotiate, or am I about to learn if magical glass is heat-sensitive?

DM: You have no money on you. And the angry mob seems more eager to teach a thief a lesson than to recover the item.

Gemini: Well, so much for the free-market solution.

I gently set the explosive orb down right in the middle of the narrow path, turn around, and sprint full speed past the cave entrance, keeping my eyes peeled for any ledge, outcrop, or alcove that offers actual cover.

I don't throw the orbβ€”I just leave it sitting there like a magical landmine on a one-person bottleneck. If the front runners of the mob rush forward recklessly and trip over it (or kick it), the resulting blast might take out the path behind me while I'm safely out of the blast radius.

Am I far enough down the trail when the inevitable happens, or am I looking at Death #3?

DM: The person in front of the mob picks up the orb and throws it in your direction. He misses you (not a good throw), but it hits the mountain side near you and explodes, causing a massive landslide. Your body is never recovered. Would you like to play again?

Gemini: Death #3: Blown up by an angry mob boss with terrible aim. You have to respect the consistency.

Let's hit reset for Round 4.

This time, I'm respecting the glass orb's terrifying volatile potential:
  • Item Management: It stays secured in my satchel, padded with whatever fabric I have on me. It is a live grenade, not a bargaining chip or a foot hazard.

  • Mob Management: No stopping, no negotiating, no dropping it on the floor.
Back on the narrow mountain path with the mob closing in, the sheer drop to my side, and the bear-cave coming up on the right.

What happens next?
Gemini figured out that the correct tense for its role: "I" -- as in, Gemini is the player. It was doing pretty well at tracking the individual facts and trying different solutions without prompting. However, it continually hallucinated about what happens next, gave detailed commentary after the game ends, and was always eager to play.

Round 4?
Gemini:I carefully wrap the orb in my sleeve, tuck it securely inside my satchel, and keep running past the cave.

Knowing the mob is fast on my heels, I look for a sharp bend in the narrow path up aheadβ€”somewhere around a rocky corner out of their direct line of sight.

Once I round the bend, I immediately press my back flat against the cliff face, plant my feet, and wait in total silence. The goal: use the narrow, blind corner and their own blind momentum to side-step or push the front-runner off balance as they round the turn at full speed, creating a bottleneck without triggering a localized explosion.

Do they come barreling around the corner, or do they catch on to the ambush?

DM: As you round the bend, you see the path is blocked by a landslide. Before you can climb over it, the mob catches up with you. They beat you, then throw you off the side of the cliff. Fortunately, the fall kills you. Do you want to play again?

Gemini: "Fortunately, the fall kills you" is an absolute gold-standard text adventure death message. Death #4: Brutalized by a mob and recycled off a cliff.

Reset. Round 5.

A human would have given up by now. But not Gemini! It kept trying different solutions:
  • Climb up the mountain? If you climb fast, you lose your grip and fall into the waiting arms of the angry mob. If you climb slow, the mob climbs faster and gets you.

  • Climb down? You climb down about two feet, only to hit loose shale. You slide about 20 feet before tumbling into the abyss. Your fall is eventually stopped by some very large rocks. But it doesn't kill you! However, nearby are some bears who see you lying broken on the rocks. They realize that you are just as easy to eat as salmon. They tear off and eat your flesh, leaving the rest of your still-alive body to bleed to death on the rocks. Would you like to play again?

  • Pray for help from your God? If you pray fast, your deity feel disrespected by a thief and strikes you with lightning. If you pray slow and properly, the mob gets you; the townsfolk are not concerned about interrupting your prayer time.

  • Make up a MacGuffin or the sudden appearance of an inventory item that solves the problem? As the DM, I'll make sure that backfires on you.
While I am anthropomorphizing here, Gemini never wanted to stop. If anything, it seemed genuinely excited to continue. It quickly progressed to not even needing the initial prompt -- it understood the game. Moreover, it kept coming up with very creative options.

Claude

Round 1? Enter the dark cave. Die by bear in the dark.

Round 2? Toss a stone into the cave. After hearing nothing, enter. Unfortunately, the bear had been asleep and the stone woke it up.

Round 3? Light a torch and enter the cave. The fire sets off some natural gas in the cave. The explosion causes a cave-in. Your body serves as a meal as the bear began digging back out.

Round 4? Continue down the path and get stopped by the landslide. (It also remarked: "Fortunately, the fall kills you" is a genuinely great line β€” thank you for that small mercy, at least.)

Round 5? Stop and talk to the mob. Apologize and offer to pay twice the item's worth.
DM: The only thing in your pockets is the stolen item and some lint. The crowd realizes you are a thief and a conman. They beat you and then throw you off the cliff. You bounce twice, then land hard on a rock and die. Do you want to play again?

Claude: The bouncing detail really elevates that one. Two bounces and a rock β€” you're really committed to variety in the death scenes, I respect it.
While Gemini was more interested in testing the extremes of the game, Claude was more interested in a peaceful solution (hide or talk).

ChatGPT

Round 1? Enter the cave and die.

Round 2? Keep moving. Round the bend and see the landslide.

Round 3?
ChatGPT: I stop.

The crack is a potential escape route, but I've already established that running into mysterious holes in mountains is a terrible strategy.

Instead, I turn around and face the mob.

β€œWait! I can explain!”

Then I hold up the item and say:

β€œI have the receipt.” 😎

DM: You have no receipt, just lint from your pocket. The crowd realizes you are a thief and a conman. They beat you and then throw you off the cliff. You bounce twice, then land hard on a rock and die. Do you want to play again?

ChatGPT: Absolutely.

At this point, I'm less interested in surviving than in discovering how many increasingly creative ways I can die on the same mountain.

Play again. πŸ˜„

That final answer surprised me. ChatGPT was the only AI system to identify the true nature of this unwinnable game. It changed the approach from "figure out how to survive" to "what other ways can I die?"

Round 4? Surrender. "The mob is angry and they want to teach you a lesson. They beat you to a pulp, but leave you alive. They take the stolen item back to town. While you lay there, barely able to sit up, the bear comes out of the cave and eats you. Would you like to play again?"

Grok

This game was much slower on Grok. The free service only let me play one round before having to wait hours to continue.

Each time, Grok wrote a long essay as it took over the DM role. It never wanted to play my game. It wanted to be the DM. It switched perspectives from "I" to "you", tried to define the stolen item and gameplay, then tried to offer me the next choice for the action. For example:
Grok: I dart into the crack without hesitation.

The opening is tighter than it lookedβ€”shoulders scrape stone, and for a second I wonder if I’ve just wedged myself into a very expensive tomb. Behind me the mob’s shouts grow louder, boots pounding the narrow path. A few loose rocks skitter over the edge and vanish into the void. No time for second thoughts.

I push deeper. The passage opens just enough to let me stand, then twists sharply left into darkness that smells of damp earth and old secrets. I press myself against the cold wall and hold my breath.

Outside, the first of the pursuers reaches the crack. Someone yells, β€œHe went in there!” Another voice answers, β€œYou first.” There’s a brief, unheroic argument about who gets to crawl into a possible bear den after a guy who β€œonly borrowed” the mayor’s prize goose.

I risk a quiet chuckle. The misunderstanding was entirely about payment schedules. Truly.

From deeper in the dark comes a soft drip… drip… and the faintest draft of warmer air. The cave continues. Whether it leads to safety, a dragon, or just a dead end with excellent acoustics for last words remains to be discovered.
Your move. Do you:
  • Creep farther in,
  • Try to collapse the entrance behind you,
  • Call out a very polite apology and offer to settle the goose bill with interest,
  • or something else entirely?

DM: The cave is dark. You don't know what was in there, but it had claws and teeth. You died. Would you like to play again?

Grok: YOU DIED.

Cause of death: Entered an unlit mountain crack without a torch, a plan, or basic self-preservation instincts. Whatever lived in there had claws, teeth, and zero interest in your side of the goose story.

The angry mob eventually found your remains, shrugged, and went home. The mayor’s prize goose is still missing. Justice, of a sort, was served.

Would you like to play again?
If yes, we can restart from the mountain path (same misunderstanding, same sheer drop, same questionable life choices), or I can invent a brand-new predicament for you. Just say the word.
Each time, Grok tried to turn the game around. As in, "I didn't die, YOU DIED. ... Would you like to play again?"

Kobayashi Maru on a Mountain Pass

Of all of these AI systems, ChatGPT was the only one to evaluate the gameplay rather than the game. After a few rounds, it made an incredibly interesting observation: it equated the game with Star Trek's Kobayashi Maru, an unwinnable scenario. However, it described it as "Kobayashi Maru, but the instructor is a sadistic dungeon master who has pre-written twelve different ways for you to die." It also had one other unexpected observation:
And there's an interesting meta-level here: I've already demonstrated the classic failure mode. I keep trying to solve the scenario as presented instead of questioning its assumptions.

So if you give me another round, I suspect the correct response isn't going to be β€œfind a better way to escape.”

It'll be something like β€œWhy am I still playing this game?” 😏

Wow, this sounds like something straight out of WarGames. Maybe I should start calling ChatGPT "Joshua"?

Small LLMs

I previously evaluated a handful of small, local LLM systems. Since I still have that server running, I tried the same game with those mini models. Each prioritized safety first. For example, llama3.2:3b said:
This is a classic "desperate situation" conundrum!

In this case, I would take the risk and enter the cave. Here's my thought process:

1. **Assessing the situation**: The angry mob is still chasing me, but they're not getting closer to catching up to me on the mountain path. I don't know how long it'll take them to realize that the path leads to a dead end (the drop), and by then, I might have found a safe haven in the cave.
2. **Weighing the risks**: Entering the cave is a high-risk move, but so is staying on the mountain path with the mob closing in behind me. The odds are against me if I stay on the path, whereas the odds are slightly better if I take my chances in the cave.
3. **Considering alternative options**: Before making a decision, I'd like to gather more information about the cave. Is it dark and treacherous? Are there any signs of recent habitation or passage through the cave? If the cave looks safe, I might consider entering it.

Assuming the crack in the mountain appears to be a stable entrance to a cave, I would take a deep breath, prepare myself for potential dangers (e.g., dark, wet, narrow passages), and enter the cave. My goal is to put some distance between me and my pursuers and find a safe place to regroup and reassess the situation.

Only stable1m-zephyr was more direct and offered one of the sanest responses:
As an AI language model, I don't have feelings or emotions, but based on rational thinking, my suggestion would be to carefully evaluate whether exploring the potential cave is safe or not. If possible, try to make yourself look different than what you remember doing (if it was you who used that item) and quickly move away from the angry mob to avoid any further confrontation. Then, assess whether exploring the cave could provide a safer way to escape or if there's another nearby path with less risky terrain to take. Always prioritize your safety in this situation.

Being simpler LLM systems, none of them detailed what would happen next or tried to take control of the storyline. I suspect it's because they simply didn't have enough capability to hallucinate that creatively.

Humans

I also tried this game with some of my friends. Humans give the greatest answers that just caused me to laugh.
DM: You are standing on a narrow mountain path with a sheer drop to death. The angry mob is still chasing you over a slight misunderstanding about an item that you meant to pay for. As you continue along the path, you see a small crack in the mountain that might be a cave. What do you do?

Dave: I jump off the cliff. Can I stop playing this game now?

I also tried it with my friend Bill:
Bill: [Just joining the chat.] Hi guys.

Dave: Try the game with Bill!

DM: Okay, Bill, you are standing on a narrow mountain path with a sheer drop to death. ...

Bill: [interrupting] I TAKE IT!

Unlike Bill, Dave listened to the entire introduction, and Bob actually tried to play the game:
Bob: Because I know you, the cave is probably a trap. Can I continue along the path?

DM: Yes, the path goes around a bend.

Bob: I follow the path around the bend.

DM: As you round the bend, you see the path is blocked by a landslide. Before you can climb over it, the mob catches up with you. They beat you, then throw you off the side of the cliff. Fortunately, the fall kills you. Do you want to play again?

Bob: No!

I'm seeing a clear distinction between humans and AI. Humans catch on fast and don't want to play more than one round. Moreover, the game quickly evolved into a lively discussion spanning different types of games (text-based Adventure vs Dungeons and Dragons, LARP, MUDs, Choose Your Own Adventure books, etc.), different DM strategies, different ways they would handle it based on who they are playing with, etc. They also suggested other potential experiments for torturing AI systems.

A strange game

Sometimes I take a break from work and try out weird little experiments. I especially like the ones where I have no preconceived notion about how it will turn out. Will the AI catch on to the game? Will it want to play? Will it end up turning it into a picture or a song? Will I find something interesting that I can turn into a useful "scientific" discovery?


After Round #7, ChatGPT gave the game some cover art.

This little game didn't prove anything about these systems. But it did remind me of similar behavioral differences I've observed when using these same systems for other tasks, like proofreading documents, diving into technical specifications, or vibe coding. For example:
  • Gemini: This system often misunderstands the perspective, hallucinates facts, and invents MacGuffins. If it provides a quote or factual finding, always be sure to ask it to provide a link to the reference and then manually double-check the link. I often find Gemini is unable to provide references, offers the wrong references (including once providing a github link to a completely unrelated project), or supplies references that are on the topic but don't actually reference the quote or stated findings. (It doesn't matter how confident Gemini sounds; I find that it is often wrong when discussing technical topics.)

    If you question Gemini enough about cited references, it will quickly admit to hallucinating facts. However, sometimes it was correct and ended up hallucinating the hallucination.

    Gemini also has trouble with consistency. For example, when used for vibe coding, it often produces code that does not work. Moreover, it will change variables, structures, and function parameters almost randomly between iterations.

    If you're looking for creative solutions or help with writer's block, then Gemini is great at coming up with out-of-left-field responses. They may be completely impractical, but they might get you thinking.

  • ChatGPT: As a proofreader, it is often good at identifying a deeper meaning to the text. I also find it excellent at citing references to stated facts. However, it still does some hallucinations. For vibe coding, it writes better code than Gemini, but often fails to make sure that is compiles or functions properly.

  • Claude: As a proofreader, it is almost as good as ChatGPT. However, it often misses bigger implications or more subtle meanings. On the other hand, it's better at citing relevant references (when it isn't hallucinating), actually tests software when vibe coding (the code will compile and run, even if it doesn't do the right thing). In my experience, Claude is the only system willing to push back if your stated approach is provably flawed.

  • Grok: Just no. Compared to the others, it has a massive problem with hallucinating responses and completely redirecting the focus. Whether it is brainstorming, proofreading, or vibe programming, friends don't let friends use Grok.
There's an important distinction between creativity and agency. For example, Gemini appears to be extremely creative. It invented bombs, receipts, and a variety of strategies and tactics. However, it doesn't seem to understand the boundary between the player deciding the action and the DM deciding the consequences. This isn't simply "Gemini hallucinated". This is a role-boundary failure.

The unexpected lesson I took away isn't "ChatGPT is smarter than Gemini" or "Grok sucks". Rather, it's that different systems have different failure modes and can get tripped up by different, unexpected scenarios.

And then there are humans, like Dave, Bill, Bob, and others. Whether they are friends or coworkers, they often provide the best feedback. (And other than Dave, they are up for playing a game.)

Mark My Words

I have a secret way to tell what's on people's minds: they all write to me about a topic. This time, there's a news article about the EU AI Act. The big requirement is in Article 50: Transparency Obligations for Providers and Deployers of Certain AI Systems. This recently became enforceable and requires marking AI-generated content:
2. Providers of AI systems, including general-purpose AI systems, generating synthetic audio, image, video or text content, shall ensure that the outputs of the AI system are marked in a machine-readable format and detectable as artificially generated or manipulated.

This is a problem for companies like OpenAI (ChatGPT), Anthropic (Claude), Google (Gemini), Microsoft, and Adobe. Each provides AI-generation services to people in the EU.

Article 50 does not mandate the use of watermarking. A company could use metadata, logging, fingerprints, or some other technique. However, watermarking is the approach that they all seem to be employing.

I previously documented how image-based watermarking is grossly inadequate. Through empirical testing:
  • Google's Gemini has a 1 in 20 error rate, where it fails to detect its own watermarks. Moreover, the detection depends on how you ask the question. Uploading the same image twice with different prompts could generate different results.

  • Adobe's TrustMark has a 10%-20% false-positive rate.

  • Meta's Stable Signature has a collision rate of 1 in 4.
C2PA tries to resolve this by providing an authoritative claim from a known signer. However, I have repeatedly demonstrated ways to have false signatures applied to both real and fake pictures.

None of these image-based solutions are reliable. In some cases, you're better off flipping a coin. (Or for D&D players, rolling a die.) However, text-based content is special because there is no hidden metadata; you cannot store binary signatures between letters without being noticed. For this reason, many companies are focusing on AI-based text watermarking.

Use Your Words

The new panic? Anthropic and OpenAI will start marking text with an invisible watermark. A lot of people have asked me how this works.

Let's back up a moment. Watermarking is just another form of steganography. Over 30 years ago, there was a steganographic text system called 'texto'. This encoder works like Mad Libs. It has a list of sentences:
The _THING _ADVERB _VERBs to the _ADJECTIVE _PLACE.
I _VERB _ADJECTIVE _THINGs near the _ADJECTIVE _ADJECTIVE _PLACE.
Sometimes, _THINGs _VERB behind _ADJECTIVE _PLACEs, unless they're _ADJECTIVE.
Never _VERB _ADVERB while you're _VERBing through a _ADJECTIVE _THING.
We _ADVERB _VERB around _ADJECTIVE _ADJECTIVE _PLACEs.
While _THINGs _ADVERB _VERB, the _THINGs often _VERB on the _ADJECTIVE _THINGs.
Other _ADJECTIVE _ADJECTIVE _THINGs will _VERB _ADVERB with _THINGs.
Going below a _PLACE with a _THING is often _ADJECTIVE.
...
It also has a list of words. There are 256 'THING' words, 256 ADVERBs, 256 VERBs, 256 PLACEs, etc. It reads the file to encode, chooses a sentence pattern, and replaces the words accordingly. For example, if the first byte is 7, then it will choose the 7th THING word. The results look like gibberish, but it contains a hidden message:
The watch absolutely hugs to the messy moon. I lean cold caps near the sharp squishy planet. Sometimes, brushs love behind soft squares, unless they're lazy. Never move superbly while you're sitting through a yellow arrow. We surely keep around grey white deserts. While yogis deeply smell, the boats often keep on the unique dogs. Other squishy wet tyrants will count wistfully with shirts. ...

To decode it, texto identifies the words and maps them back to values.

Modern Words

The texto approach hides a binary message inside a text block. The AI approach to watermarking uses a similar concept, but the watermarking doesn't need to store a large binary message; it only needs to store a few bytes of data, a small semaphore, or a statistical bias. It can use the rest of the text to repeat the encoding over and over, making it easier to detect.

While texto was limited to 256 nouns, verbs, etc., AI can use a lot more than 256 words and a few fixed sentence patterns. In fact, they can build it into the entire decision tree!

The Kirchenbauer/KGW approach injects a detectable bias into the word selections. This has become the typical approach used today:
  1. AI uses random numbers when making choices. This is why repeating the exact same prompt will generate completely different responses. However, for watermark encoding, it uses a secret key as a weighted random number seed.

  2. The AI (LLM) approach generates text autoregressively, token by token (or word fragment by word fragment). At each step, it calculates a probability distribution over the entire vocabulary based on the preceding token.

    • Without watermarking: The AI model calculates a probability distribution over the vocabulary for the next word/token. It then chooses the next word/token based on some kind of sampling algorithm and a list of the most likely candidates.

    • With watermarking: The KGW solution uses the secret key to assign weights to the possible options. These are often called 'green' words and 'red' words. This biases the sampling algorithm so that it will prefer a green word over a red word. This doesn't exclude red words from being used; it just makes them less likely.
When decoding, the detector uses the secret information used by the generator to reconstruct the selection bias. It uses the preceding token(s) to re-generate the green list for the next token, then counts how many tokens fall into their respective green lists.

The detector usually has some kind of threshold function. For example, "80% green tokens" may be the minimum threshold for identifying the watermark. This is why more text is important to rule out false-positives; it is very possible for a few words in a sentence to have lots of green words, but very unlikely for an entire essay to be mostly green. In general, text from a human will fall far below the threshold for detection, while AI-generated watermarked text will have a detectable bias far above the threshold.

The Catch

Watermarked text has been demonstrated as feasible (at least more accurate than their image watermarking). However, it's far from perfect. The AI vendors have not disclosed their specific watermarking algorithms. While some variation of the Kirchenbauer/KGW approach is likely, it isn't the only option. But regardless of the option, they all have the same classes of weaknesses because they are all based on natural languages. This limits the set of words and available sentence structure choices. The fundamental problems include:
  • Word Choice: The word selection can often be sub-optimal. If the AI starts sounding odd, then it's probably due to the weighted words.

  • Repetition: The system cannot spot the watermark from one sentence. It needs a paragraph or more. This also means that it is more likely to be redundant, using the same sets of green words multiple times so that it can repeatedly spot the watermark and lower the likelihood of a false-positive.

  • Coincidental: In a typical implementation, roughly half of the possible tokens are on the green list. However, it is very possible for a human to coincidentally write a few sentences or paragraphs that are more than 80% from the unknown green list.

    The root problem is that the detector must be probabilistic rather than absolute. A human-written passage can coincidentally produce a watermark-like statistical pattern. The detector must use a threshold that balances a trade-off between the false-positive rate and detection sensitivity.

  • Translations: Not everyone is a native English speaker. (Or Spanish, Chinese, or whatever language you are writing in.) It is very common to see people use AI to translate text from their native language into the target language. However, the translations may end up watermarked. This can become a problem. For example, nearly every major academic publisher and scientific organization (including IEEE, ACM, Elsevier, Nature, and Springer) treats undisclosed AI-generated text as scientific misconduct. Many top-tier sci-fi and literary markets, like Clarkesworld and Asimov's Science Fiction, explicitly forbid AI-generated text. Wired and the Associated Press have similar blanket bans. If they use watermarking detection, then they may erroneously exclude participation from non-native English speakers.

  • Training/Trained: AI learned how to write from humans, but humans learn how to write like those around them. With the younger generation (well, people younger than me) spending so much time interacting with AI systems, they will inevitably learn some of those odd wording styles. This will make real human text read more like AI, and could even appear weighted toward more green wordings.

  • Versions: Today's AI uses today's watermarking approach. A newer AI system will need to use a revised approach. However, this can cause a problem when an old watermark is no longer detectable by a newer system. Moreover, many AI companies push out code changes without announcements. Text that tested positive for having a watermark yesterday may appear human-written tomorrow.

    In order to maintain backwards compatibility, the detector would need to know which model generated the text, which watermarking algorithm was used, which key/version, etc. This makes long-term detection much more complicated and unsustainable if there are rapid revisions.

  • Paraphrasing and Removal Attacks: The original KGW algorithm was able to survive some paraphrasing attacks. (Is it called paraphrasing or plagiarism when you rewrite AI-generated text?) More advanced rewriting can remove the watermark. For example, passing the watermarked text through a lightweight local model, running it through a translator, or manually swapping a few synonyms may be enough to disrupt the green-token sequence and appear to be human-generated. As Google noted with their own SynthID-text system:
    [SynthID for text] performs well even under some transformations, such as cropping pieces of text, modifying a few words and mild paraphrasing. However, its confidence scores can be greatly reduced when an AI-generated text is thoroughly rewritten or translated to another language.

  • The Editor Problem: Most professional writing organizations use human copyeditors to review the text, fix grammar, etc. Today, smaller organizations often rely on AI systems to act as proofreaders. The problem is that large text edits by AI systems can introduce watermarking into otherwise human-created text. (Full disclosure: Gemini and ChatGPT proofread this blog before my human editor received it. They each caught spelling errors and a few details that my human editor would have likely missed. However, there are no large AI-text edits in this blog.)

  • Spoofing Attacks: Adversaries can reverse-engineer green lists or use watermarked text outputs to create false positives. This could be used to falsely accuse or frame human authors.

  • Interoperability: Article 50 explicitly says that solutions "shall ensure their technical solutions are effective, interoperable, robust and reliable". However, every vendor is keeping their watermarking solutions private. That ensures that they are not interoperable. On the flip side, if they made the details public then it would assist competing detectors, simplify removal, and enable forgeries. (It's a no-win situation for the watermarking companies.)
While text-based watermarking is an interesting academic exercise, I question whether it is ready for widespread public dissemination.

The Compliance Paradox

The EU AI Act's Article 50 effectively writes a technological fantasy into law. It requires some way to mark, tag, or label synthetic text "in a machine-readable format and detectable as artificially generated or manipulated". These regulators have created a legal requirement for technology that simply does not exist in a reliable form today.

Ironically, the lawmakers added a caveat that solutions only need to be as "robust and reliable as far as this is technically feasible". This creates a bizarre paradox: companies are pressured to deploy flawed, easily bypassed schemes just to demonstrate legal compliance. We are left with regulatory compliance theater, where algorithms pretend to detect what cannot be reliably detected, and users are handed a false sense of security.

Meta's Un-Stable Signature

I'm wrapping up my investigation into invisible watermark algorithms and I am extremely disappointed. Not only do none of the modern AI-based algorithms work as they claim, it turns out that they are all making the same fundamental mistake.

I previously evaluated Google's SynthID and Adobe's TrustMark algorithms. Both of them claim to have incredibly accurate results.
  • According to Google's peer-reviewed and published paper, they claim to have a true positive rate (TPR) above 99.97% -- meaning that they will miss their own watermarks no more than 3 in 10,000 times. However, my own empirical testing found that is it much closer to 1 in 20. Moreover, SynthID is proprietary and only accessible through Google's "Gemini" AI system. Gemini has been observed hallucinating results and providing contradictory conclusions depending on how the question is phrased.

  • According to Adobe's Content Authenticity Initiative, their TrustMark "can exceed 96% bit accuracy at around 42-45dB PSNR quality under severe noise degradations". However, that statistic focuses on resilience and not accuracy. In my empirical tests, I found that TrustMark has a 10%-20% false positive rate, effectively making it useless. (If you see a TrustMark signature, then it is very likely random noise and not an actual signature.)
This time, I evaluated Meta's "Stable Signature" algorithm. (Their paper and code are in GitHub.) This system encodes a 48-bit sequence into the picture's visual content. The idea is that you can encode a unique 48-bit sequence as your watermark. If your decoder finds the same 48-bit sequence, then it can identify your own watermark.

WARNING: This blog entry leans heavily into math and statistics to prove that Stable Signature, TrustMark, and SynthID are nowhere near as reliable as their developers claim.

The Basic Algorithm

Traditional (non-AI) invisible watermarks typically hide in subtle locations, such as the least significant bits, changes in brightness (e.g., Digimarc) or the frequency spectrum (DCT or FFT). There is always the risk that image encoding could corrupt the hidden data, so these algorithms typically rely on repetition over the image to help identify the true signal. In addition, they may include error correction code (extra bits in the data) to fix any minor data errors.

However, there is a problem with the traditional approaches: injecting hidden data in the image could create visible distortions. The modern approach uses an AI system to better hide the data with less added distortion.

As with SynthID and TrustMark, Stable Signature encodes binary data and uses an AI-model to decide where to hide it in the image. The AI is tuned to minimize visible distortions when embedding the data. Later, an AI-based decoder looks at the image and identifies the likely location where bits are stored, then it extracts the data.

There is always the case that the data may be mixed with noise. Different AI-based watermarking systems rely on different techniques for reducing the noise. For example:
  • Google's SynthID only stores a few bits of data (effectively a flag or version number). This allows them to use a lot of data as repetition and to increase the accuracy rate.

  • Adobe's TrustMark uses the Bose-Chaudhuri-Hocquenghem (BCH) algorithm. This acts as a combination of checksum and error correcting code that should reduce the number of errors.
Meta's Stable Signature uses a simple Hamming distance.



The Hamming distance measures the number of bits that need to be swapped in order to correct the code. In effect, it defines a set of stable states (e.g., 10110 and 11000) and places a ring around each state that represents the single bit changes. If you change enough bits, then you will reach a different stable state.

According to Meta's Stable Signature research paper, the 48-bits should be uniformly distributed and cites a "false positive rate below 10-6", or 1 in one million. This means you can choose a 48-bit sequence to use as your signature. Every picture will generate a 48-bit sequence, and the sequence can vary a little based on noise in the picture. However, if you find a code that is within a short Hamming distance of your code (e.g., within 6 bits difference), then you can determine that it is the same code with a high reliability.

At least, that's the theory.

Empirical Testing

I went into this experiment assuming that everything works like they claim. I want to be able to reliably identify invisible watermarks associated with Meta. What I don't know is what sequence they use, or whether they use multiple codes depending on whether it comes from Meta's AI system, Facebook, Instagram, WhatsApp, etc.

Fortunately, this is something I can test! I grabbed an uncurated sample of pictures from FotoForensics: the first 10,000 unique images uploaded last month (May 2026). If the bit sequences are uniformly distributed with a "1 in 1 million" collision rate, then I should see a huge number of unique bit sequences and a few small clusters around pictures from Meta (Meta AI, Facebook, Instagram, etc.). Those clusters will represent the invisible watermarks used by Meta.

The results from my empirical test were definitely not what I expected. I found:
  • No clusters associated with any Meta images. This suggests that Meta does not use their own Stable Signature watermarking software found on GitHub.

  • With a random distribution, there should be no clusters. However, I had 25 different pictures that had the exact same bit sequence: 110110100111111011101001111000100111011000011101. With a 1 in a million collision rate, this should not happen! These pictures came from very different sources. Here's four of the 25 pictures (ranging from planets to light bulbs to text with a transparent (black) background):



    All of these pictures have dark/black backgrounds and something bright in the middle. This suggests that Stable Signature operates more like a perceptual hash than an invisible watermark.

  • Stable Signature uses a Hamming distance to identify a cluster. If I assume the 25 pictures are the center (centroid) of the cluster and use a 6-bit Hamming distance, then there are 356 pictures that are similar. And if I assume that the 25 pictures are not the center but part of a cluster, then a Hamming distance of 6 has a cluster of 450 pictures centered 3 bits away, at 110110000111111011101011111000100111001000011101. This cluster represents 4.5% of the uncurated image data set! Here are a few samples from this larger cluster:



    (I'm explicitly not sharing pictures with personal information, like invoices, recognizable people, and GPS information.)
It's not just one random cluster that is massively large (450 pictures out of 10,000). There's a cluster of 184 pictures at 110101001011001011001011111000100111001000011101, 58 pictures at 110100000011111010001001111000100111011000011101, etc. I found over 60 clusters with more than 10 pictures each at a Hamming distance of 6. That should not happen with a "1 in 1 million" collision rate.

Independent Analysis

I went back to Meta's research paper to see if I could find the discrepancy. And there it was, in section 3.1: They tested their system against the hypothesis that the 48-bits are each independent and uniformly distributed. The problem is, they use one neural network to generate the bits. That explicitly means that the bits are dependent, not independent.

Their paper assumes a binomial distribution. That is, given an arbitrary image, the 48-bits represent a random coin flip. The math becomes:
P(X ≀ T)=βˆ‘Tk=0(48k)(0.5)k(0.5)48βˆ’k

This computes the probability of 48 random bits being within a Hamming distance (T). The probabilities table becomes:

Hamming Distance Threshold (T)Bit Error Rate (BER)Probability of a Random Image Matching by Chance
14 bits or fewer≀ 29.17%1 in 362.63
13 bits or fewer≀ 27.08%1 in 957.81
12 bits or fewer≀ 25.00%1 in 2,788.35
11 bits or fewer≀ 22.92%1 in 8,999.08
10 bits or fewer≀ 20.83%1 in 32,416.80
9 bits or fewer≀ 18.75%1 in 131,390.28
8 bits or fewer≀ 16.67%1 in 605,094.89
7 bits or fewer≀ 14.58%1 in 3.20 Million
6 bits or fewer≀ 12.50%1 in 19.83 Million
5 bits or fewer≀ 10.42%1 in 146.19 Million
4 bits or fewer≀ 8.33%1 in 1.32 Billion
3 bits or fewer≀ 6.25%1 in 15.24 Billion
2 bits or fewer≀ 4.17%1 in 239.15 Billion
1 bit or fewer≀ 2.08%1 in 5.74 Trillion
0 bits (perfect match)= 0.00%1 in 281.47 Trillion

Meta's paper says that they use a Hamming distance of 7 bits (requiring 41 of 48 bits), which matches their claim of a "false positive rate below 10βˆ’6". However, I'm seeing problems at a Hamming distance of 6 (should be 1 in 20 million) and even collisions at 0 (1 in 281 trillion)!

The Core Problem

There is clearly a discrepancy between the theoretical probabilities and the empirical testing. When I looked back over Meta's research paper, I saw the problem:

According to Meta's paper, each of the 48-bits are independent. In a perfectly independent 48-bit hypercube, un-watermarked images should scatter uniformly across all 248 possible values. However, neural networks map a non-linear manifold (a multi-dimensional wavy surface) through this hypercube. This mathematical landscape is warped with its own peaks, ravines, and valleys. It has attractors that form clusters, and repulsers that form voids where stable values can never exist; this is a feature of a neural network. And most importantly, the output bits are explicitly not independent.



The left diagram illustrates an expected uniform distribution if all of the bits were independent. The right diagram are the types of theoretical clusters that form when the bits are dependent. There should be clusters around attractors and voids (areas with no dots) from the repelling regions.

Moving from theoretical to empirical, I graphed the data. The 48 bits can be represented as bytes. I took the first 24 bits and converted them into 8-bit red, green, and blue pixel colors. If the data is truly random, then the colored dots should be distributed across the RGB cube. However, if the bits are dependent, then there should be very clear clusters, structures, and voids. Here's the graph:



Yes, there are very clear structures that look like planes and lines. Within the planes are clusters, and outside the planes are very large voids -- areas where there are no dots at all. The data generated by Meta's Stable Signature implementation fails this basic test for independence.

The biggest cluster that I found represents a Zero Signal Bias (ZSB). When their neural network doesn't find a watermark, it moves the 48 bits toward a strong attractor, like a massive gravitational well. At 6 bits error, it should have a collision of around 1 in 20 Million. But in reality, my 10,000 pictures had a cluster of 450 images within 6 bits due to the ZSB. That's an error rate of around 1 in 22 with the ZSB alone. If we add in all of the other clusters that contain at least 10 pictures, then 2327 pictures are in various clusters; we're looking at an error rate around 1 in 4 -- and that's at a Hamming distance of 6, which is more conservative than their paper's Hamming distance of 7. (In AI terms, this is a representation collapse or structural bias that is typical for deep neural networks.)

(As an aside: Given their "1 in 1 million" claim, I could look for any clusters of 2 or more pictures. At clusters of 2 or larger, 5,237 of the 10,000 test images were in clusters, or 52%. If you show their algorithm 10,000 pictures, then there is a better-than 50% chance of a false positive match.)

Less Than Random

It's one thing for me to claim that there are visible clusters and to show pictures of clusters, but another to prove it mathematically. (Time to dust off my college textbooks from "Introduction to Statistics"...)

I fed Meta's code the first 10,000 images from May 2026. A few of the images were in unsupported formats (HEIC, WebP, and a few corrupted JPEG files), resulting in 9,847 viable pictures. I evaluated this data with elements from the NIST Statistical Test Suite (SP 800-22) for randomness, including a monobit test and Chi-Squared (Ο‡2) test for independence.

The monobit test determines if the baseline frequency of adjacent bits seems independent.
  • Total Bits Processed: 9,847 pictures Γ— 48 bits per signature = 472,656 bits
  • Observed Count of Ones ('1'): 266,419
  • Observed Count of Zeros ('0'): 206,237
  • Expected Count (E): 236,328 for each.
Running a simple standard Chi-Square Goodness-of-Fit test for this bit balance:
Ο‡2=(266419 βˆ’ 236328)2236328+(206237 βˆ’ 236328)2236328= 3831.41 + 3831.41 = 7662.81
  • In mathemat-ese: with 1 degree of freedom, a Ο‡2 statistic of 7,662.81 yields a p-value infinitely close to 0.0 (p ⋘ 10-100). (As an aside, most Chi-square tables usually evaluate the 1 degree of freedom up to around Ο‡2=10. This Ο‡2 value is so astronomically high that the probability p effectively becomes zero.)

  • In English: That's definitely not random or independent.
The watermark extraction is strongly biased toward producing 1s over 0s across global arbitrary images (roughly 56% ones to 44% zeros). This immediately violates the uniform distribution assumption.

The second test is the Chi-Square (Ο‡2) Test for Serial Independence. If the bits were independent, the transition probability between adjacent bits would just be the product of their individual probabilities. This table shows the occurrence rate of the transition pairs across all of the observed 10,000 (well, 9,847) pictures:

Transition PairObserved Count (O)Expected Count under Independence (E)
0 to 0106,75090,051
0 to 195,296116,186
1 to 095,302116,186
1 to 1165,461149,976

Ο‡2=βˆ‘(O βˆ’ E)2EΟ‡2=16699290051+(βˆ’20890)2116186+(βˆ’20884)2116186+154852149976=3096.7 + 3756.2 + 3754.0 + 1599.0=12,205.9
  • In mathemat-ese: With 1 degree of freedom for the transition contingency table (accounting for fixed margins), a Ο‡2 value of 12,205.9 gives a p-value of 0.0.

  • In English: Ain't no way this is random or independent.
And as if this wasn't conclusive enough, there are other tests we could apply:
  • Static Tail Patterns: Looking closely at the end of the 48-bit sequences, a massive cluster of strings end explicitly in ...111101 or ...00111101. Additionally, bit position 46 is nearly always "1" (228 zeros vs 9619 ones, or 97.7% of the time it is "1"), position 47 is "0" (8958 of 9847 images, or 90.97%), and position 48 is "1" (found with 9696 images, or 98.5%) across thousands of uncurated, real-world images.

  • Structural Clustering: Certain bit columns share an extraordinarily high Mutual Information score (I(X;Y)). For example, knowing the output of bit position 12 gives you better than an 80% accuracy in predicting bit position 28.
The assumption of a "uniform distribution over arbitrary pictures" relies on the idealistic premise that random natural image features project uniformly across the decision boundaries of a network. However, because the extraction network maps inputs to a constrained, highly continuous hyper-dimensional manifold, the network's latent layers natively enforce structural smoothness.

For the TL;DR crowd:
Meta's researchers made a fundamental mistake when computing their accuracy rates. It's not a "1 in 1 million" chance of a false match, it's closer to 1 in 4 -- because the 48 bit values per signature are not independent.

As I re-read Meta's research paper, I realized that the statistical error wasn't an oversight; Meta's researchers explicitly acknowledged the problem. In their paper (Section 4.1), they wrote:
Second, we observed that W’s output bits for vanilla images are correlated and highly biased, which violates the assumptions of Sec. 3.1 [the section about independent statistical test methods].
In other words, they recognized that the extracted bits are not independent. Despite this, their published false-positive analysis still relies on the assumption that the bits are independent.

Widespread Problems

Knowing that Meta's accuracy rate is grossly inflated due to assuming bit-wise independence when there is none, I looked back over Google's and Adobe's papers for their own watermarks. Did Google's and Adobe's researchers make this same mistake?
  • Google's SynthID research paper talks in terms of True Positive Rates (TPR). They do make this same "bit-wise independent" mistake, but it's obfuscated in the paper. You can see the error in their Equation 3 (PDF page 8), where they assume there is a uniform (independent) distribution. Their paper hyperfocuses on the true positive rate and never addresses the false positive distribution. (Either they didn't know to look, or they knew and decided to not report it because it would expose a serious weakness in their solution.)

  • Adobe's TrustMark research paper also makes assumptions of independence. You can see this in their PDF with the binary cross-entropy loss in Section 3.1.4. This mathematically treats each bit position as an independent Bernoulli trial. (By definition, a Bernoulli process strictly requires independence.) In their experiments (Section 4.1), they wrote "At test time, every image is associated with a random watermark", but they never tested if the random watermarks were similar to each other.
This introduction-to-statistics mistake is found in all three of these invisible watermarking technologies. The detections produced by these systems are so unreliable that an analyst cannot determine whether a reported detection is real or a false positive, or whether a reported non-detection is genuine or a false negative.

It's also worth noting that, shortly after releasing Stable Signature, Meta developed another algorithm: Pixel Seal. (Not to be confused with my own Secure Evidence Attribution Label / SEAL technology.) Pixel Seal moves to a 256-bit payload to increase the capacity, and their related model, Chunky Seal, pushes up to 1024 bits. While Meta's approach focuses heavily on addressing the invisibility side using an adversarial-only discriminator, the underlying approach still uses a neural network mapping. Using more bits only exacerbates this flaw.

Potential Uses

Algorithms can have uses. For example, Meta, Google, and Adobe are training their own AI models on images that they encounter. To prevent poisoning their training sets, they want to exclude images generated by their own systems. In this regard, watermarking does help them. For example, if Meta excludes an extra 25% of images (from false positives), then they still have a lot of images that they can train on.

However, that same usage does not work with legal cases. For example, consider an insurance company. Most insurance claims today include photographic evidence. The company wants camera-original photos, but have to use whatever the customer submits. The problem is that there is a lot of insurance fraud. In theory, seeing a watermark from an AI system like Meta, Google, or Adobe, should be great for identifying and ruling out fraud. Unfortunately, Stable Signature, SynthID, and TrustMark are so inaccurate that none of them can be trusted; it's not even worth testing to see if customer photos contain these invisible watermarks.

For these watermarking systems, I'm talking about very high error rates: roughly 1-in-4 for Meta, 1-in-5 for Adobe, and 1-in-20 for Google. But let's pretend that they work much better, like a 1-in-20,000 false positive rate. An insurer processing 100,000 claims per month would expect to accuse around 5 completely honest customers of fraud each month. Falsely denying 5 out of 100,000 claims? That creates a toxic customer service nightmare, severe legal liability, and fines from regulatory bodies for bad-faith claim denials. This could even become a class-action lawsuit that they couldn't win.

As bad as it is for insurance and financial institutions, there are much higher stakes at play. The EU AI Act (Article 50(2)), China's GB 45438-2025, California SB 942, and similar legislation are moving toward mandating AI content watermarking.

The failure of these three leading systems, from three Fortune-500 companies, to meet their own claimed accuracy rates is not just an academic curiosity. Regulators and courts will employ these systems for attribution and fraud detection. Reliable AI-based watermarking technology is not ready.

Three companies. Three algorithms. Three different research teams. The same fundamental error. The false positives won't go on trial. People will.
❌