Anthropic on Friday published a blog post aimed at answering basic questions about how to watermark text generated by its chatbot Claude. How do watermarks actually work? Can they be hidden by editing? And how does this affect the code?
Claude users have been discussing the move since the company revealed earlier this week that it was doing the watermarking to comply with transparency provisions in EU AI law that require AI companies to use systems that allow them to identify AI-generated content.
For example, on Reddit, one poster characterized this as a conspiracy against innocent Claude users, while another claimed that “the only reason not to want this is to lie to people.” And Business Insider reports that “dozens” of X users claim to have canceled their Claude subscriptions as a result.
Anthropic’s new post begins with an overview of the concept of watermarks, explaining that when making “low-stakes choices,” such as choosing between “cloudy” and “gray” to describe the weather, Claude can create patterns in a response that are “undetectable to the reader, but detectable to someone who has the key to encode it.”
“Watermarks do not affect the quality of Claude’s output,” the company said. “For the reader, a watermarked response is indistinguishable from a non-watermarked response.”
Specifically, Anthropic said it plans to release a watermark detection API using the SynthID-Text approach outlined by the Google DeepMind team in 2024. He also noted that watermarking is different from the AI detection approaches offered by companies like Pangram, which look for “content” in text (such as the syntax “He’s not (X), it’s (Y)”) to reveal AI usage. “Detecting these patterns is fundamentally different from checking a watermark.”
Could someone rewrite the text to hide the watermark? Anthropic said it was possible, but that “a simple edit would not completely remove the watermark,” but “a complete rewrite that replaces every word could.”
“Of course, in the latter case, it is debatable whether the text can no longer be said to have been generated by AI,” the company said.
As for whether a watermark is detectable in text that was proofread or edited only by Claude, Antropic said it depends on “the length of the text and how much editing was done by Claude.” If only lightly edited, “nearly every word” would have been written by a human author, and “very little (if any) would be watermarked.”
On the other hand, code should have less watermarks than other text. This is because the model requires you to write code to work, and you don’t have the freedom to choose from a variety of equally valid options.
“That said, watermarks can be used in areas where you can arbitrarily select specific words or terms in your code, such as comments in your code,” Antropic said. “But by definition, it has negligible impact on the actual code that is generated.”
Anthropic also said Claude is not the only AI chatbot to generate watermarked text, as “other major model developers have signed the same code of practice and will be implementing their own watermarks.”
If you buy through links in our articles, we may earn a small commission. This does not affect editorial independence.
