Skip to content

Anthropic shares more details about how Claude’s new watermarks will work

Anthropic publisheda blog postFriday seeking to answer some basic questions about how it will watermark the text generated by its chatbot Claude. Such as: How will the watermarking actually work? Can it be hidden with editing? And how does this affect code?

Claude users have been debating the move since the companyrevealed earlier this weekthat it would be doing this watermarking to comply with the EU AI Act’s Transparency Code, which requires AI companies to use systems that make it possible to identify AI-generated content.

On Reddit, for example, one postercharacterized this as a conspiracy against innocent Claude users, while another claimed, “The only reason you wouldn’t want this is to lie to people.” AndBusiness Insider reportsthat “dozens” of users on X have claimed to cancel their Claude subscriptions as a result.

Anthropic’s new post starts with a general overview of the watermarking concept, explaining that when making “low-stakes choices” — like choosing between the words “overcast” and “grey” to describe the weather — Claude can create a pattern in its responses that is “undetectable to the reader, butisdetectable to anyone who has a key that encodes it.”

“Watermarking does not impact the quality of Claude’s output,” the company said. “To a reader, a watermarked response is indistinguishable from an unwatermarked one.”

More specifically, Anthropic said it will be using the SynthID-Text approach thatthe Google DeepMind team outlined in 2024, and that it plans to release a watermark detection API. It also noted that watermarking is distinct from the AI detection approaches offered bycompanies like Pangramthat look for “tells” in the writing (like the construction “his isn’t [X], it’s [Y]”) to reveal AI usage: “Picking up on these patterns is fundamentally different from checking for a watermark.”

Could someone just rewrite the text to hide the watermark? Anthropic said it’s possible, but “light editing probably won’t remove the watermark completely,” while “a complete rewrite where every word is replaced will.”

“In the latter case, of course, it’s arguable whether the text can any longer be described as AI-generated,” the company said.

As for whether the watermark will be detectable in text that was only proofread or edited by Claude, Anthropic said that will depend on “the length of the text and how heavily Claude has edited it.” If it’s only been lightly edited, “nearly all the words” will have been written by the human author and “there’s very little (if anything) for the watermark to attach to.”

Code, meanwhile, should have less of a watermark than other text, because the model will need to create working code and won’t have the freedom to choose between a variety of equally valid options.

“Having said that, in areas where there is an arbitrary choice between particular words or terms within the code, the watermark can be used, such as comments within code,” Anthropic said. “But by definition, it will have a negligible effect on the actual code produced.”

Anthropic also said that Claude won’t be the only AI chatbot to generate watermarked text, as “other major model developers have signed the same Code of Practice and will be implementing their own watermarks.”

Topics

When you purchase through links in our articles,we may earn a small commission. This doesn’t affect our editorial independence.

Anthony Ha is TechCrunch’s weekend editor. Previously, he worked as a tech reporter at Adweek, a senior editor at VentureBeat, a local government reporter at the Hollister Free Lance, and vice president of content at a VC firm. He lives in New York City.

You can contact or verify outreach from Anthony by emailinganthony.ha@techcrunch.com.

Scale faster. Grow your portfolio. Gain practical expertise. No matter your goal, Disrupt can empower you.Save up to $300 today!

  • If Apple sends you a push notification alerting you to a spyware attack, take it seriouslyZack Whittaker

If Apple sends you a push notification alerting you to a spyware attack, take it seriously

If Apple sends you a push notification alerting you to a spyware attack, take it seriously

  • Zack Whittaker
  • Some Claude users are mad that Anthropic’s new watermarks will catch them using it at their jobs, classesLucas Ropek

Some Claude users are mad that Anthropic’s new watermarks will catch them using it at their jobs, classes

Some Claude users are mad that Anthropic’s new watermarks will catch them using it at their jobs, classes

  • Lucas Ropek
  • Delta investigating after someone set up fake Wi-Fi network mid-flightLorenzo Franceschi-Bicchierai

Delta investigating after someone set up fake Wi-Fi network mid-flight

Delta investigating after someone set up fake Wi-Fi network mid-flight

  • Lorenzo Franceschi-Bicchierai
  • Anthropic says it will watermark text generated by its AI modelsIvan Mehta

Anthropic says it will watermark text generated by its AI models

Anthropic says it will watermark text generated by its AI models

  • Ivan Mehta
  • Mark Zuckerberg’s AI manifesto is exactly why people don’t like AIRussell Brandom

Mark Zuckerberg’s AI manifesto is exactly why people don’t like AI

Mark Zuckerberg’s AI manifesto is exactly why people don’t like AI

  • Russell Brandom
  • YouTube now requires creators to have twice as many watch hours to start earning moneyAisha Malik

YouTube now requires creators to have twice as many watch hours to start earning money

YouTube now requires creators to have twice as many watch hours to start earning money

  • Aisha Malik
  • This ‘adversarial’ pattern can prevent surveillance cameras from detecting youZack Whittaker

This ‘adversarial’ pattern can prevent surveillance cameras from detecting you

This ‘adversarial’ pattern can prevent surveillance cameras from detecting you

  • Zack Whittaker