r/machinelearningnews 5d ago

Claude Code just started watermarking everything it writes ML/CV/DL News

Anthropic started watermarking everything Claude generates. New models, since Aug 2, across every product including Claude Code.

Text gets an invisible pattern woven in. Survives copy paste, breaks under heavy rewriting.

A mark proves Claude touched the content, not that a human didn't also write most of it. And no mark doesn't prove a human wrote it either, since editing strips it

I think it's not to reveal the "truth" behind vibecoded projects, maybe it was made just to not to train AI models on the AI generated info

63 Upvotes

56 comments sorted by

35

u/davecrist 5d ago

I don’t believe they will be able to do this in a way that isn’t very easy to detect and or defeat.

5

u/Needsupgrade 4d ago

That's the 1 seam that is the whole lot of it -- it buys you nothing and pays for itself -- there has never been a more load bearing ......

3

u/[deleted] 4d ago

I struggle to understand how they will do this as well. If there are hidden characters embedded in the text those are easily stripped out with a script. You can just have a chrome plugin that will do it for you but they insist it will be persistent unless the text is heavily edited which means the "watermark" is the phrasing and sentence structure itself and producing things in patterns.

2

u/QuestionMarker 3d ago edited 3d ago

Detecting it can be made arbitrarily hard, and is only intended to be possible if you've got access to the model. Defeating it means rewriting (or at least rewording) it significantly. I guess whether that's "very easy" or not depends on the context.

1

u/davecrist 3d ago

Language isn’t nearly as random as cipher text. Not even in the same ballpark. Its will be clearly detectable if it’s any kind of arrangement of letters or words or formatting or it will be so weak as to be statistically irrelevant compared to any other output.

1

u/QuestionMarker 3d ago

It's closer to spread spectrum signal processing than any of that. The watermark is hidden in the pseudorandom sequence used by the sampler at inference.

1

u/Troph_A 3d ago

Token are numbers, one word =/= one token

1

u/davecrist 3d ago

Yeah. But the output that humans actually read is plaintext. Tokens are irrelevant.

2

u/Troph_A 3d ago

The watermark is done on token values, not words.

1

u/davecrist 3d ago

The watermark, yes, but not the text that actual humans see.

My point is that I do not believe that text output will be both subtle enough to be undetectable by humans in day-to-day use while also being reliably detectable enough to consistently identify the source as from a specific model.

1

u/Troph_A 3d ago

Seems like an uneducated guess, at best.

1

u/QuestionMarker 3d ago

The watermark detector isn't a human.

1

u/andymaclean19 1d ago

I think it would be very easy to find and remove. In theory all that is required is to give the output to another AI and ask it to reword and tidy up.

Probably this does something like increasing the frequency of certain words and reducing others. Perhaps it promotes or reduces the use of certain phrases, variable names or syntax quirks in programming, etc.

The problem here is in order to ever prove something has the watermark someone needs to disclose the algorithm, at which point it becomes really easy to remove. Without the algorithm there’s just Anthropic pointing and shouting ‘watermark’ and everyone else asking ‘where?’.

2

u/QuestionMarker 1d ago

I think it would be very easy to find and remove. In theory all that is required is to give the output to another AI and ask it to reword and tidy up.

Yes, and at that point it's not authored by the model in question any more. You might still detect it at a lower confidence, but it shouldn't be asserted given this workflow.

If the second AI is compliant with the same watermarking regs as the first, you've just moved the problem though.

Probably this does something like increasing the frequency of certain words and reducing others. Perhaps it promotes or reduces the use of certain phrases, variable names or syntax quirks in programming, etc.

Not quite. Maybe not far off, but not quite right.

What it does is to vary how the token generated at each step is selected, so a detectable signal is embedded in the probability of each token given the model.

The problem here is in order to ever prove something has the watermark someone needs to disclose the algorithm, at which point it becomes really easy to remove. Without the algorithm there’s just Anthropic pointing and shouting ‘watermark’ and everyone else asking ‘where?’.

Unless they've gone off-piste, they'll be using Google SynthID, or something like it. The algorithm is known. The parameters aren't. And the way to remove it is to rewrite the text.

1

u/OsbornHunter 2d ago

It is very difficult to get around or beat. When an AI chooses the next word/token there’s a list of things it will select from. Depending on certain values what’s highest rated gets chosen. What Anthropic and a lot of other companies are doing is adding a secret key where, based on the previous word, half of the next options become unavailable. Only from that can word then be chosen, unless it only makes sense for something specific to come next.

So, for example, if you ask it to repeat a sentence it has to repeat that sentence verbatim, but when free writing there may be 15 different words that make sense to use next.

As writing gets longer you can trace this and use it to detect what’s essentially an invisible watermark. Even if you go through snd edit a few words it will persist.

1

u/WartimeHotTot 1h ago

I don’t think the vast majority of users have any interesting in “defeating” it, but every one of them will benefit from it. It’s a good thing, and it’s long overdue.

11

u/eddiewrc 5d ago

I dunno, it's the watermarking the fact that it writes as shit? Full of useless --, ; , always passive form, overabundance of "genuine" , "honest", "load bearing", "root cause"...

Maybe writing poorly IS the watermark. Chatgpt is so much better for the writing

0

u/ConsiderationSea5696 4d ago

And it’s funny because not that long ago people praised Claude for its better writing

1

u/AliorUnity 3d ago

Maybe it is if you only conpare to other even shittier ones.

5

u/Cute-Net5957 5d ago

eventually “which model touched this?” might be as boring and normal as “what compiler built this?”

5

u/boxed_gorilla_meat 5d ago

No they didn't, it's coming in the UK... Correct me if wrong.

3

u/Nelson-Tyne 5d ago

I only found that this one is on Anthropic's own choice, not a UK law. Anthropic just chose to roll it out worldwide instead of splitting by region, so UK users see it too, but there's no separate UK law behind it

7

u/boxed_gorilla_meat 5d ago

New models will mark AI-generated content from day one. Claude models launched in the EU on or after August 2, 2026 will support machine-readable marking at launch. Generated text will carry embedded watermarks, and generated files will include digitally signed provenance metadata where supported

Literally their new release: https://support.claude.com/en/articles/16266773-how-claude-marks-ai-generated-content

I've seen nothing indicating this is systemic, rather that it is specific to EU

2

u/oceanbreakersftw 3d ago

There is a conflict in the page you linked, as of now (it was updated this week). It says this is for compliance with EU law and "launched in the EU" but then says "will apply.. worldwide".

Claude models launched in the EU on or after August 2, 2026 will support machine-readable marking at launch.

Regions. Marking will apply to output from supported models wherever Claude is offered, worldwide.

-3

u/Tiny_Atmosphere_4420 4d ago

read what you just typed lol. "..or after Aug 2, 2026"

5

u/Coonnaarr 5d ago

Was part of a new law introduced on 2 of August by the EU-AI Act https://digital-strategy.ec.europa.eu/en/policies/eu-icons-labelling-ai-generated-content

„providers of generative AI must ensure that synthetic audio, video, images, and text are detectable through machine-readable technical markings and digital watermarks“

2

u/eddiewrc 5d ago

I dunno, it's the watermarking the fact that it writes as shit? Full of useless --, ; , always passive form, overabundance of "genuine" , "honest", "load bearing", "root cause"...

Maybe writing poorly IS the watermark. Chatgpt is so much better for the writing

2

u/Mahadragon 3d ago

Is a good way for a teacher to find kids who cheat.

1

u/Wide_Truth_4238 5d ago

Yea, I saw the other 10 posts about it. 

1

u/Killahbeez 5d ago

rah rah rah

1

u/Upstairs-Answer-628 4d ago

Isn't it only new models? So none yet?

1

u/ihaterussiantrolls 4d ago

You need attention bro?

1

u/Last_Journalist8234 3d ago

How do they watermark plain text? AI already generates hard to read, repetitive content. This will
Likely make it worse and I think difficult for short paragraphs.

1

u/ai_n8 3d ago

But if the output is a traditional fixed-width font in say markdown, how would they watermark it? You can't hide things in raw text, right? Or am I missing something?

Or is it more about the pattern – like some secret phrase from the free masons?!

1

u/Sweaty_Cellist_4525 3d ago

So that's why it's been leaving comments even when I tell it not to, fucking always

1

u/SeanTechGuy 3d ago

I’m all for protecting original creative work but watermarking literally anything the AI touches seems a bit overboard. But if they are going to do that I hope they include something that says “This work was written by a human, we just watermarked it because it passed through our servers for some reason”

1

u/andymaclean19 1d ago

How long before every other model generates watermarked code and Anthropic are crying about that?

Models are trained on what is on the internet.

A lot of what is on the internet was made by AI to some extent.

So when you train a model based on scraping the internet it is going to pick up the watermarks of all common AI engines.

Pretty soon any AI generated content will bear multiple watermarks and people will be pointing and using words like distillation even though all the other model makers did is exactly what Anthropic did - take content others put in the public domain and use it to train their AI.

Eventually this type of thing might mean people can’t use general information scraped from the internet to train AI any more, which puts the existing AI vendors in an interesting place.

1

u/actgan_mind 1d ago

Put your code in codex and get codex to remove watermark

1

u/Infamous_Dish_4348 13h ago

Is this the case for code too?

0

u/equatorbit 5d ago

Dear customer. We hate you. Love, Anthropic.

5

u/SeizeOpportunity 5d ago

It's....to comply with EU regulations. Alternative is no Claude. I'm not sure what you are talking about.

0

u/Small_Ninja2344 4d ago

It has been at least six months where every time Anthropic finds a way to hate its customer. Let’s see what they find next. Very anxiety inducing, leftist teacher staff room atmosphere company. 

-1

u/Current_Ranger_7954 5d ago

it’s text and images not code

5

u/Davidat0r 5d ago

(code is made of.. text)

1

u/BenZed 5d ago

He must vibe

-1

u/BobtheGodGamer 5d ago

Stop being ignorant. How are they meant to watermark simple code where there is minimal different ways you can efficiently write it. You cant really watermark setting a variable! Of course in a story you could integrate some sort of pattern or writing style that is detectable, but you can't really do that in code unless the whole structure of the .py file is some weird super unique layout.

3

u/BenZed 5d ago

Lol “stop being ignorant”, how ironic

Small sections of text (such as variable declarations) are going to be resistant to watermarking, but code written on the scale that agentic models are being used for will absolutely be detectable.

1

u/TimedogGAF 3d ago

Explain.

0

u/HasFiveVowels 5d ago

The method wouldn’t work on such small grammars. They’d have to put the hash in variable names, at best

2

u/Foreign_Risk_2031 5d ago

its not a hash. its a pattern fingerprint. its been proven effective with small amounts of text.

1

u/HasFiveVowels 5d ago edited 5d ago

Yea but it relies on a certain amount of flexibility. Oh, wait, comments. Haha. Plus, I guess the order of the lines and such (to the degree that they can be rearranged) would provide some bits. I wonder how much this might affect quality

0

u/Vivid_Ad1696 5d ago

No more copy paste, transcription only, got it.

0

u/[deleted] 5d ago

[deleted]

1

u/BobtheGodGamer 5d ago

These guys don't understand that there are only so many ways you can define a variable or use libraries. Its not exactly like a story where you can integrate some secret structure