Lies And Scams Taint Watermark Removal Apps Now That Anthropic Started Watermarking Claude AI Outputs
Anthropic's new watermarking of AI-generated text from Claude has triggered a rush for watermark removal apps. These watermarks are subtly derived using statistical word choices, making them hard for humans to discern, but can be detectable by the right tools. However, the article notes that these watermarks can be easily compromised through editing, integration into larger texts, or re-processing by other AIs. A major concern is the surge of fraudulent removal apps exploiting this demand, often making false claims or even containing malware. This problem is anticipated to worsen as more AI providers adopt watermarking, creating a fertile environment for scams and deception, thus users are urged to exercise extreme caution.
In today’s column, I examine the flurry of lies and scams underlying those watermark removal apps that are supposed to be able to remove digital watermarks found in the outputs of generative AI and large language models (LLMs). Though such lies and scams have been around for quite a while, they have taken on a new and egregious life after Anthropic announced that it is watermarking the AI-generated output from Claude.
Let’s talk about it. This analysis of AI breakthroughs is part of my ongoing Forbes column coverage of the latest in AI, including identifying and explaining key AI complexities (see the link here).
The place to start is by discussing Anthropic’s recent announcement about watermarking. It has been like the blast of a starting gun for a slew of unintended adverse consequences, as you’ll see in a moment.
In a posting on the Anthropic Claude support page on August 11, 2026, the popular AI maker announced that they are starting to watermark their AI outputs. For my in-depth analysis of this matter, see the link here. The upshot is that when you use Claude to answer questions or provide responses to your prompts, the plain text that is generated will henceforth contain a secret watermark. The text will look perfectly normal. Nothing obvious to the naked eye can discern that the text has been watermarked. There aren’t catchy emojis or oddball characters being implanted.
How can they possibly hide a watermark in ordinary text and yet you cannot see it? Aha, this is cleverness in mathematics and computational orchestration to select words that ultimately have a subtle but detectable statistical pattern. When the AI is composing a response, it is carefully selecting words that not only answer your question or query but also reflect patterned choices of which words to use, acting as a non-obvious signal of sorts.
Read More: Dodgers Sign Chadwick Tromp to Bolster Catching Depth on Minor League Deal.
Think of it this way. Suppose that any given sentence can be composed of words that have multiple choices of which word to use in the sentence. For example, a sentence might say that a cat sat on the floor. Another way to say that same sentence is to indicate that a feline resided on the ground. Assume that those words, such as feline for cat, reside for the word sat, and the ground for the floor, are all second choices, yet are still fully reasonable choices. The algorithm inside the AI is choosing the words that embody a pattern, such as always picking the second choices of word selections, that can later be detected.
I think you can see that this statistical uplift is going to be quite hard to detect. Humans are unlikely to see the watermark by looking for any patterns in the wording. All the sentences are still going to make sense and abide by whatever the topic at hand is. The subtlety of picking the second statistically viable word on numerous occasions is a nearly hidden way of producing the watermark.
How does an authorized detection tool figure out if the watermark is present?
Aha, that’s by knowing what approach was used at the get-go while the text was being watermarked. The chances of any usual detection method ferreting out the watermark are low. A tool that is built knowing the specific method can examine the sentences and compare the word choices to the pattern of word choices that the AI would normally make. If the second word choice is consistently being encountered in the examined text, this is a strong indicator that the AI indeed generated that content.
We can make this method much more robust. Maybe instead of always choosing the second choice, the watermark process does something else. Suppose that 50% of the time the second choice is made, 30% of the time the third choice is made, and 20% of the time the fourth choice is made. This makes things even harder for anyone else to crack and find the watermark. An even stronger method includes having a secret cryptographic key that guides the watermarking process toward the preferred token patterns.
I’ve so far been explaining how text-oriented watermarking takes place. The statistical uplift scheme is one of many mathematical and computational methods that can be utilized. Much simpler approaches can be used, but those are typically readily defeated without much effort involved. If the text contains emojis or special characters as watermarks, you will undoubtedly remove those visible disturbances without hesitation. There might be so-called invisible characters too, such as using a white font on a white background. Again, that is trivial to find and expunge.
What's Your Reaction?
Like
0
Dislike
0
Love
0
Funny
0
Wow
0
Sad
0
Angry
0
Comments (0)