Rendered at 19:37:24 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
Dilettante_ 11 hours ago [-]
I notice the "Limitations" section talks about how content only at some point touched by Claude may return a positive, and content that returns a negative may still be Claude generated. But I really would have liked for them to state explicitly that entirely false positives where a piece is fully human-written may still be marked as generated, because too many institutions with the power to ruin someone's life over that have trouble understanding the concept.
dragonwriter 38 minutes ago [-]
> But I really would have liked for them to state explicitly that entirely false positives where a piece is fully human-written may still be marked as generated, because too many institutions with the power to ruin someone's life over that have trouble understanding the concept.
This is marketing material aimed, in part, at encouraging the usage you are concerned about, which is why they do not highlight that problem.
unprovable 8 hours ago [-]
This. FN rates are cute, but FP rates will ruin an academic career or a student's work/further study choices if their content gets marked erroneously. Surely the answer is a sequence of marks?
Keen to see if they are doing something SynthID-esque?
WD-42 5 hours ago [-]
Do they care about false positives? As long as it’s even somewhat reliable that’s enough for them to prevent training on their own slop. I think this is a big reason to do this that’s overlooked.
DennisP 4 hours ago [-]
Good point. But if that were their only purpose, there'd be no need to share it with anybody. In fact, they'd get the best results by not mentioning it.
WD-42 4 hours ago [-]
That’s true. But they probably want to be able to identify other models slop as well. And with the laws popping up, it makes sense to do it the way they are.
jrflo 4 hours ago [-]
I think that false positives are inevitable due to the method of watermarking being embedded in the text itself. The output is intended to mimic human writing, therefore it's entirely conceivable that a human could by chance write text that contains the watermark. The odds may be extremely small, but it's not something you could ever guarantee.
andai 3 hours ago [-]
I keep hearing how humans are thinking and writing more and more like AI.
I think in this case I think it's some kind of cryptographic signature smeared across the token IDs, so I don't think the risk is very high.
jrflo 32 minutes ago [-]
I read the original paper they're basing this off of and I think you're right. I do wonder how much of a quality tradeoff there is with perturbing the next token probability distribution. My intuition tells me that a more "prominent" watermark will necessarily degrade output quality. If they are trying to balance quality and watermark prominence, I wonder if that affects the FPR.
xena 53 minutes ago [-]
You're absolutely right! Humans have been slowly thinking and writing more and more like AI. As people get more and more exposed to the stochastic patterns of large language model tools, it's normal for them to emulate the styles of communication they are exposed to. This is commonly called "brainrot" by those in Gen Z and younger cohorts.
If you find yourself getting to be afflicted by this "brainrot", be sure to go outside and take a moment to ponder what's around you. The grass is there and will be there long after we are all gone. Consider this for a moment as your organic thought processing unit starts to slowly munch away at its internal context window.
ed_elliott_asc 2 hours ago [-]
It’s worse than that, false positives are possible but someone generating text should be able to get ai to change some words and formatting to break the watermarking, then ai detectors can tell them how well they did.
I don’t know what the answer but I absolutely know it isn’t this.
dbqpdb 55 minutes ago [-]
I think we need a chain of custody system for content, but that would require browsers, software, websites, operating systems, phones, camera manufacturers, etc to all get on board. But each intermediary or source (optionally) cryptographicaly signs a piece of content that it either generates, edits, or passes along, and the end result at a destination, is that content is either 'trusted' if its cryptographic chain is solid, or un-trusted otherwise.
dragonwriter 35 minutes ago [-]
> I think we need a chain of custody system for content, but that would require browsers, software, websites, operating systems, phones, camera manufacturers, etc to all get on board.
It would also require the individual humans you are trying to control to get on board otherwise the analog hole breaks the chain, absent mindboggling levels of physical surveillance on top of the the total monitoring of all electronic data flows that this idea requires.
akersten 7 hours ago [-]
If there is any false positive rate (which, because text will naturally and by chance include tokens from the green and red sets in some pattern, there will be), tools making promises like "detect AI-generated text" are unacceptable. They are going to turn innocent people into pariahs on some unsubstantiated "this content is 37% likely to be AI" claim that the user has no way of verifying or inspecting more deeply, we just have to trust the statistical box and assign some meaning to whatever that number means. 37% of my phrases are AI? There's a 37% chance my entire text is AI written? Part of the fun is not knowing!
This is scripture homeopathy and it's irresponsible.
wrsh07 7 hours ago [-]
I'm curious about your thoughts on pangram. I only really see posts on Reddit claiming it falsely labels their content as ai generated but nobody will actually post examples of "textbook from twenty years ago" or upload screenshots of a journal (also those posts usually feel deeply ai generated without an ai detector)
Do you think this is an impossible task and we shouldn't try to solve it? Or do you think it's doable and that some ai detectors might be better than others?
nemomarx 6 hours ago [-]
This feels testable - you could go to fanfiction or similar sites with billions of words of writing from before 2016 or so and run them through it.
I tried a chapter just now and got human doing that, but I'm not invested enough to run a hundred samples today. But it sounds like it would be an alright way to audit it? I will confess I'm pretty skeptical you could ever eliminate false positives here though. I can often get an ai sense from some writing on my own but I doubt it would be better than 90% accurate, and "ai plus human editing" might screw with that anyway, stuff like that. I would have preferred we just never developed this kind of thing so I wouldn't have to guess.
philote 6 hours ago [-]
That still might work better with older texts. As AI-generated text gets more prevalent, I'm guessing people will start subconsciously adopting AI writing styles.
nemomarx 6 hours ago [-]
Yeah, that's one of my questions. Everyone who talks to AI for too long seems to get worse at writing anyway, and humans mirror any form of conversation to some extent.
subsistence234 2 hours ago [-]
LLMS aren't the only thing that has changed over time in the way texts are written.
if they used older texts as training data, to some extent pangram would just be an age classifier for writing style.
StilesCrisis 6 hours ago [-]
It's been done and showed up on HN recently. Older content was quite consistently marked as not-AI.
estebarb 5 hours ago [-]
Language distribution shifts. Eventually people will start adopting the distribution used by LLMs, making classification harder.
Also, this doesn't even consider the case where people use LLMs to translate their original works. Or people that use it for spelling/grammar checks.
Personally, I believe these checkers do more harm than good. Any false positive can ruin someones life.
> Pangram 4 achieves a 0.0041% false positive rate (roughly 1 in 24,000) on 1,000,000 human-written English FineWeb evaluation examples
> Overall False Negative Rate is 0.3396% on English AI generations (26 generator models)
nunez 6 hours ago [-]
pangram is pretty good; i use it all of the time and pay for it. surprised that it's not mentioned that often here. they just released a new model that is supposed to lower the fpr (false positive rate) even further than it's already impossibly low score. it also detects AI in images now, though I expect the fpr to be pretty high there given its newness.
whack 5 hours ago [-]
> But I really would have liked for them to state explicitly that entirely false positives where a piece is fully human-written may still be marked as generated
Are you talking about pieces that were fully human-written with zero AI editing/rewriting etc? If so, what makes you think that false positives will happen there? They aren't looking for "writing styles" or emdashes etc. They are using watermarks and metadata.
If you're talking about people using AI to copy-edit text they manually wrote, this was explicitly called out in the article:
> A detected mark provides a signal that content was processed by Claude, but is not fully conclusive. Detecting a Claude mark tells you that the content may have been processed by Claude. It does not, on its own, confirm the full provenance of the content. For example: Claude may not be the original author. People often use Claude to proofread, translate, summarize, or convert files. The output can carry a Claude mark even if the underlying ideas, text, or data originated from another source; The content may have changed after Claude processed it. Marked content may be modified, excerpted, or combined with other material after Claude processed it.
Dilettante_ 3 hours ago [-]
The former. I'm not sure what you mean by metadata, but my expectation was that anything that Claude could put into the plaintext to identify itself may plausibly also accidentally be produced by [a million monkeys on typewriters/one in a million human writers], since in the end, the writing is using the same language and symbols that humans use. How unique could the LLM possibly make it while still retaining its usefulness?
WiSaGaN 8 hours ago [-]
My guess is that they will later "reveal" some "violations" but provide little evidence citing proprietary algorithm.
m00dy 4 hours ago [-]
Deepwalker once cracked Gemini's watermarking system, I'm sure they will also work on this [0].
If LLM training data is human-written, and LLM output mimics that input, how could you not have false positives?
SkyBelow 5 hours ago [-]
Because it won't be in the training directly. It is applied after a model generates its distribution of likely tokens, biasing each token randomly based on a random key and unrelated to any meaning of the words. So half the time, the most likely token becomes more likely and half the time it becomes less likely, and the same for every other token (when temperature is above 0).
You then look at the tokens actually picked to see how closely they follow this pattern that isn't connected to the meaning of the tokens. With enough text, you can then analyze the chance of it happening by chance verses being because the generation of the tokens was done using the algorithm, and you can save a positive result until you are arbitrarily sure. There is a chance of a false positive, but the chance of a false positive approaches the chance that the murderer happened to have fingerprints that matched your and both forensics labs happened to have mixed up the dna tests and the eye witness happened to misremember the face and your phone gps happened to glitch out and put you at the murder scene at the time of the crime all happening. It is theoretically possible only in the same sense that quantum teleporting a cat is theoretically possible.
The real question is how much text do they need for a given level of certainty and what do they check for. If they flag a positive at a p value <.01, that's a problem. If they can reasonably get a p value of < 1e-12 in only a few paragraphs of text, that is effectively no false positives (but a lot of 'too short to analyze' outcomes).
embedding-shape 10 hours ago [-]
Read said section yourself perhaps.
wrsh07 7 hours ago [-]
Somewhat trivially, if I ask Claude to transcribe an image and then check if that transcription is ai generated it will likely say yes.
Many users are not smart enough to realize that the transcription step is where the ai (watermarks) were necessarily injected.
basch 6 hours ago [-]
How is a perfect transcription of an image watermarked?
wrsh07 22 minutes ago [-]
It depends on how it does watermarking!!
Note, there are many ways to represent words visually on computers that look identical
FeteCommuniste 4 hours ago [-]
"Perfect as far as human perception can tell" is a weaker standard than "bit-to-bit copy." Maybe it's that?
suddenlybananas 10 hours ago [-]
That's essentially impossible, unless you mean they didn't measure a false positive rate.
Filligree 9 hours ago [-]
For watermarked long-form text, it is actually possible. Makes the watermark more fragile, but the math is considerably more forgiving than usual.
embedding-shape 9 hours ago [-]
> For watermarked long-form text
What about "watermarked long-form code"? I'm having a hard time understanding how a model could watermark not prose, but functional/semantic text like code, that actually has meaning. You can't switch our the characters, you can't use various types of whitespace, you can't add arbitrary code comments, and a lot of other restrictions. Is there any state of the art methods for watermarking code without affecting the quality/correctness?
wpietri 9 hours ago [-]
As anybody who has put together a coding standard knows, there are a lot of options for individual expression, meaning a lot of room for things like watermarking. And of course you can add arbitrary comments; my Claude-generated code is very verbose.
embedding-shape 9 hours ago [-]
> there are a lot of options for individual expression, meaning a lot of room for things like watermarking
The way I use LLMs (and I'd advice everyone to do the same) there really isn't, the agent implements things exactly how I want them, or I use the agent to massage it into the exact bit-by-bit version I imagined when I first sent the prompt afterwards. I honestly don't know what the point would be to let the agents compose worse code than what I'd do manually, although I know it's a popular approach taken by many.
> And of course you can add arbitrary comments; my Claude-generated code is very verbose.
So watermarking for all users who allow code comments from agents, no watermarking for us who force the agents to never write a single code comment? Alright, I'd be fine with that.
wpietri 8 hours ago [-]
From what I've seen, your approach to LLMs is exceedingly rare, so I suspect it's one the people who care about watermarking aren't very concerned with.
And the reason to let Claude make worse code than a professional would by hand is basically suppressed demand. Since programmers are expensive, previously code mostly got written when a large number of dollars were on the line, or when an individual programmer did something not economically optimum (e.g., hobby project).
That left a whole lot of somewhat less valuable software unwritten. It's the economic space that no-code tools have been nibbling on for years. One way to think of things like Claude Code is as effectively no-code tools. Pre-LLM no-code tools would produce data structures that got executed by special environments without ever being seen or tuned by a human. Claude Code can be used just like that, with text as the input and python as the intermediate representation that nobody ever looks at.
That approach probably isn't sustainable for what we professional programmers would call a serious project. Claude can easily get in over its head and I expect that its code decays over time, in a fashion similar to how many human teams get in a state where they just have to rewrite everything. But faster, I'd expect.
But there are a lot of unserious projects that previously would have never been created. E.g., a quick app to manage your little league team, or a bit of in-house business stuff in the "a little hard to do with a spreadsheet" range.
olmo23 8 hours ago [-]
> I honestly don't know what the point would be to let the agents compose worse code than what I'd do manually
You never generate throwaway code used to test an external service? or try out an interface idea? There's a lot of code that's only meant to be ran once. I often dont even care what language it's written in.
embedding-shape 8 hours ago [-]
> You never generate throwaway code used to test an external service? or try out an interface idea?
And save/persist it? No, most of any experimental stuff goes into /tmp which gets cleared out on reboot, nothing I care to save in any repository. Or just "show me how this would look like" and then it's only in the session itself (and the logs/state I suppose, technically...).
SkyBelow 5 hours ago [-]
For straight generated code it'll likely need more text, but it'll still show up.
In cases where one token is extremely likely, it'll randomly be red or green and still be picked in either case as it is simply the best (or only) option. So you'll have more tokens that don't show a pattern either way (half of these cases will match and half won't, just the same as if a human wrote it). Meaning you'll need more instances where multiple tokens were all likely to see if there is a pattern. Given the check algorithm can't identify these cases, it can only judge on the overall text, so the more strict a language, the more the length requirement scales.
Where I wonder if this keeps working is in tool calls. Often, you don't take code straight from the llm, you take the results of a tool call to edit already existing code. It might be that the result of this leads to far too few signals to pick up, meaning that this only works when one does significant generation with a single model (even swapping between different models, at least by different companies, breaks this just as much as having a human write parts of the code).
Think of it like finding a loaded dice. A dice that has a slight bias in a few dozen roles is just random chance. If that bias continues after hundreds of thousands of roles, the dice is loaded. But will a code base have enough samples, especially when edits made from tool calls? I could see this being unable to detect things at the size of a reasonable PR and only being useful for massive sets of changes and only if the person behind them didn't structure their AI usage to avoid detection.
peyton 9 hours ago [-]
[dead]
OGWhales 7 hours ago [-]
And yet, it remains possible that a human could write the same sequence of characters.
no_multitudes 3 hours ago [-]
How often do you add seemingly-random zero-width unicode characters to the text you write?
bufbupa 6 hours ago [-]
Sorry you're getting downvoted, this interpretation doesn't seem that far fetched to me.
Here's the strawman: The text-based watermarking is going to be done procedurally instead of generatively. Maybe they add some sequence of zero-width Unicode characters to all generated text at certain intervals. Then, there is effectively no false positive possible (because humans would [effectively] never type such sequences of unicode naturally). It may survive some editing (depending on how you select/edit the characters), and it's possible to be stripped (false negatives).
wtfwhateven 6 hours ago [-]
Why would you say something so ridiculous?
jobigoud 2 hours ago [-]
I think they mean it like this: imagine you ask me a random number sequence. I give you a random number sequence. Little did you know, I used a very specific PRNG to generate it, so later I can prove with certainty that your number was generated by me, and you can't say you came up with it yourself.
There is no room for false positive here in the same way you can't randomly find a collision in a hash function if it's strong enough. Like the rate is so infinitesimal that it is effectively zero.
Now replace random number sequence with prompted string of words. And instead of using the PRNG on every word I use it every n words. If the generated text is sufficiently long I can tell by matching the expected deterministic pattern.
You can defeat it by changing the words yourself and triggering a false negative but there isn't really any room for a false positive if the text is long enough and the pattern matches perfectly. If the pattern doesn't match then I can compute a probability.
nunez 6 hours ago [-]
I think Pangram is way ahead of Anthropic on this with their custom dataset.
simonw 21 hours ago [-]
> When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won’t see it, and it doesn’t change the meaning, quality, or readability of Claude’s response.
I'd like to know a lot more about how that works.
A lot of my interactions with Claude return pretty precise text. If I ask it to edit a project and refactor a specific function in several places I know exactly what I want to happen, it will NOT be OK if those refactors have some kind of weird pattern baked into their text to act as a watermark.
I guess this may be covered by this:
> Content generated by Claude may not carry a detectable mark if, for example: [...] The passage is very short, leaving too little text for a reliable signal;
wrsh07 6 hours ago [-]
Scott Aaronson talks about his project at OpenAI to do this^
You can carefully select which pseudorandom number generator (prng) you use to be able to id text of a certain length. I expect there is some performance characteristic you have to manage since you're doing this on every inference, but once you do that it doesn't change the output in any meaningful way (the prng is still a statistically valid prng, it just happens to let you check if the output used that prng)
The point is that you can do this simply by swapping to a different RNG, which isn't noticeable to the end user, and while it changes the output, it's not any different from how using a different seed or being lumped in a different batch will change the output.
^ excerpt:
> So then to watermark, instead of selecting the next token randomly, the idea will be to select it pseudorandomly, using a cryptographic pseudorandom function, whose key is known only to OpenAI. That won’t make any detectable difference to the end user, assuming the end user can’t distinguish the pseudorandom numbers from truly random ones. But now you can choose a pseudorandom function that secretly biases a certain score—a sum over a certain function g evaluated at each n-gram (sequence of n consecutive tokens), for some small n—which score you can also compute if you know the key for this pseudorandom function.
teravor 40 minutes ago [-]
this would be a very heavy watermark application.
there are many simpler methods, for example you can have a tiny windowed transformer operating on the output text and all you do is alter certain words (that don't change meanings) to maximize its surprise. the tiny language model will have a special training regime to build up a somewhat unique view of the language.
we are talking about a 0.5 bit watermark here (existence). I would have zero confidence in being able to reliably remove such a watermark from pretty much any medium.
wrsh07 25 minutes ago [-]
That's actually much worse because it fundamentally changes the output, whereas this doesn't change the output, it just changed the prng
That's what I excerpted, although I had seen it presented from his talk at Stony Brook
denverllc 3 hours ago [-]
I’ve always wondered how this works when we only observe the final output and not the internal state that’s used to generate the output.
The LLM presumably generates f(input, RNG) but we only can observe f(RNG).
Groxx 2 hours ago [-]
Since they do have the input, they could probably just store checksums at each step...
... though I'm not sure why that would be preferable over a coarse rolling checksum over all of the output. Seems like that wouldn't influence output, would be equally imperceptible, and probably easier to calculate (compared to "hash seed times running all LLMs supported times number of RNG algorithms, to see if output matches").
Presumably there's some other trick, or it's a red herring / failed experiment and not what they actually do in practice.
wrsh07 26 minutes ago [-]
No you don't need to do that, the prng is detectable if you know what bias to look for and have the key
COAGULOPATH 15 hours ago [-]
>I'd like to know a lot more about how that works.
My guess is that it works like Gemini's SynthID: by altering the logprobs of the next token.
Like, for every 10th token, instead of outputting the most probable, it outputs the 17th most probable, or something. (Obviously it's way more complicated but I think conceptually this is how it works.) No human will notice this, but a classifier trained on Claude's output will.
So it's not like the watermark is the words "le epic bacon" and Claude will output "le epic bacon" in everything. That would be extremely annoying (and easy to defeat).
staticman2 7 hours ago [-]
> Like, for every 10th token, instead of outputting the most probable, it outputs the 17th most probable, or something.
Wouldn't you need the prompt to know the probability of the next token?
Chabsff 6 hours ago [-]
Not necessarily. Here's a rough example (it's not what's going on here, just a representative idea):
There are words/tokens that are heavily correlated to the prompt (a yes or a no, for example), and then there are others that are going to be less so (adjectives with a lot of synonyms for example).
Given a text, you can identify what the "load bearing" and auxiliary words/chunks are. Then, looking only at the auxiliary words/chunks, you should, in principle, be able to determine what other wordings could have gone there instead. From this, you can, very roughly, recreate the token probability distribution that was in effect when those tokens were generated. With the probability distribution in hand for enough chunks of text, you can start inferring properties about the RNG process that was used to sample from those distributions.
But then, this notion of "load bearing" vs "auxiliary" can be expressed directly in the probability distributions. A load bearing token just has a very high probability, and thus any RNG bias that may have been in effect will likely be swallowed in the distribution. So the parts of the text that are highly dependant on the prompt will naturally not be contributing much information about he RNG in the first place.
JohnMakin 4 hours ago [-]
So this is finally what "load bearing seam" means.
thunfischtoast 13 hours ago [-]
They still need to choose when to do that though. When I prompt the program to e.g. alter a bash script in a specific way or to recite a longer known text it can't go round and randomly exchange tokens. It has to somehow define what is a simple repeated text from a different origin and what is a novel generation.
user43928 12 hours ago [-]
I am wondering how that applies to newly generated code.
Odd variable naming? Stylistic choices that are watermarked?
Or as someone else noted further down in the comments, it could be more subtle:
Between the first and second most likely choice, in certain positions it will consistently choose in a certain way.
melvinroest 11 hours ago [-]
> Odd variable naming? Stylistic choices that are watermarked?
Whatever it is, I'm sure it's load-bearing.
asdfsa32 11 hours ago [-]
You're absolutely right. But it is not just load-bearing, it is the load-bearing seams.
silversmith 6 hours ago [-]
Personal observation: Opus 5, over the last week, has started outputting A LOT more comments. Despite my global instructions being full of variations on "don't use comments unless absolutely necessary".
I might be imagining things of course. But comments would be great fit for this use case.
StilesCrisis 6 hours ago [-]
They have mentioned that their system prompt used to say "avoid over-commenting" and it no longer does. They should bring that back IMO.
pram 3 hours ago [-]
Yes the length of comments Opus 5 leaves is exhausting. Not to mention it will insert info thats only relevant within the current session. I've just been deleting all of them lol
kuboble 8 hours ago [-]
I cannot imagine the code with well defined specification will have extra watermarks unless the watermark is requested as part of the harness instructions.
If it works like people describe - on the every nth token or something - then the mark will be left in the chain of thought and discussion with the model - not in the code artifacts.
wpietri 9 hours ago [-]
I would guess they're not worrying about watermarking a tweak to a human-written program. That's both a tiny fraction of Claude use and of very little concern to the kinds of people who want to check watermarks.
__MatrixMan__ 6 hours ago [-]
If you're that targeted with your edits, then do you deserve a watermark anyway?
WinstonSmith84 4 hours ago [-]
it's going to be easy to defeat either way, like SynthID is. From apps built to remove the watermark, to simply rephrase the work with an Open Model...
Less probable also means less optimal and you get a subpar response. More so if it's baked into its reasoning. It's intelligence will suffer unless this is some post processing thing.
shawnz 7 hours ago [-]
There is already some intentional randomness in token selection, because it actually improves the quality of responses if you intentionally don't always pick the most likely next token.
You can hide data in that randomness without impacting the quality of the response by using a sufficiently "random looking" pseudorandom bit stream instead of real random numbers.
Well, if the model that page uses also falls under the EU act, the output will just be watermarked differently. ;)
arcfour 7 hours ago [-]
> Honest note: Anthropic has not shipped a public Claude watermark detector yet. This tool uses rewrite-based neutralization — a meaning-preserving paraphrase with a non-Claude model — which is the attack path watermark research points to. Not affiliated with Anthropic.
Well, they should have run their own AI slop website through their tool...
nunez 6 hours ago [-]
This attack was actually pointed out in the watermarking paper linked above. The researchers added an instruction to the prompt that switches letters like a Caesar Cipher. It lowers the quality of the output from the LLM but alters the "red list" enough for a watermark detection tool to fail at detecting the watermark.
fl0id 4 hours ago [-]
also their example for rewriting just completely changes it. might as well redo it in this case (with another model or by hand)
cush 4 hours ago [-]
Should be ensloppifier.app - it somehow makes the AI sound more like AI, while also completely changing the meaning and context of the input text
From their before/after:
- Certainly! -> (removed)
- onboarding redesign -> revamping the introductory process
- this week -> (removed)
- empty states -> empty sections
- CTA heirarchy -> call-to-action sequence
- interviews -> discussions
- aligned copy with brand voice -> verbal identity
... these choices change the meaning of the text
shinryuu 12 hours ago [-]
Though if pangram should be trusted, there are still statistical artifacts that tells you that a text LLM generated. I don't find that to be implausible.
timpera 10 hours ago [-]
Alas, Pangram should not be trusted.
s_dev 9 hours ago [-]
So was it going down:
"Neutralize engine is temporarily unavailable. Try again."
infinite_spin 10 hours ago [-]
My guess is it will be similar to how Genius watermarked lyrics, using things like variants of punctuation
That was actually the cause of an issue I had a couple of years ago: I had hand-typed JSON using my iPad into GitHub’s online text editor and Safari helpfully used “pretentious quotes” instead of "old-school quotes" - and the JSON library used by the program to read that file had relaxed parsing rules that accepted actual JS object literals without quoted property names; so the fancy-quotes were interpreted as part of the key-names. This took ages to figure out because when human-eyeballing the JSON file it looked perfectly fine in Notepad.
kccqzy 2 hours ago [-]
I don’t doubt your experience, but many people are intimately aware of the use of proper Unicode quote characters, in any reasonable font they choose. For me, one of the first things I learned when using LaTeX is how the quotes are transformed from the source to the typeset document; since then I’ve become extremely sensitive to the kind of quotes I see.
miohtama 13 hours ago [-]
Maybe there is a reason why Opus 5 produces such word salad conversations
andrewgleave 6 hours ago [-]
Yes. The irritating epigram / aphorism style it now uses is such a regression compared to previous Anthropic models. Probably is the case that this is due to watermarking - though hardly subtle if it is.
myko 8 hours ago [-]
So frustrating to use. And the comments generated by Claude today are unreadable garbage.
cush 4 hours ago [-]
Certainly it can't watermark text with low entropy. If you're renaming a function using claude it won't be marked
w_for_wumbo 21 hours ago [-]
What happens if someone handwrites a Claude output, then someone uses that handwritten text as a reference.
Now you've got a watermarked idea which may have no direct linkage to the usage of Claude.
TheOtherHobbes 9 hours ago [-]
If the algos work as advertised, watermarked token sequences have an extremely low probability. Copying the words by hand doesn't change that.
The mechanism seems to survive editing. The extreme probabilities get a little less extreme, but are still extreme enough to be distinctive.
But it wouldn't survive paraphrasing, because the output would be entirely human and the token correlations would disappear.
It might not survive referencing if only a sentence or two is used.
The practical issue is how true the claims are. It's one thing to create a proof of concept, another to see how it works in use.
And this is potentially catastrophic for code, because the grammar and word choices of code are completely different and more fragile than standard English.
phainopepla2 21 hours ago [-]
How is that different from referencing digital text that someone copied and pasted from Claude?
w_for_wumbo 17 hours ago [-]
Because there's an expectation of authenticity from the written word.
If you've referenced something handwritten, you don't expect it to be the output of an LLM.
Similarly, if you quote someone word-for-word, you wouldn't anticipate their words to be flagged as Claude content, but if someone memorized Claude output word-for-word. That would still be classified as a Claude output.
Going forward you could categorize the influence of Claude on a population based off a percentage match between their spoken words with the LLM prose.
dns_snek 12 hours ago [-]
Are you worried about being accused of using LLMs to generate your work? As long as you don't plagiarize you have nothing to worry about.
Cthulhu_ 11 hours ago [-]
I'm not too sure about that, people making stuff have already gotten penalized by overzealous AI detectors, most recently Kurtzgesagt.
platinumrad 11 hours ago [-]
You can't make a blanket statement like this without knowing how the watermark is implemented.
AlecSchueler 12 hours ago [-]
What if I unknowingly read content written by Claude in various articles and it influences my own writing style?
stabbles 21 hours ago [-]
It will just thread some load-bearing seams through the paragraphs.
mihaelm 21 hours ago [-]
> have some kind of weird pattern baked into their text to act as a watermark.
public abstract class BaseAnimalBeanFactoryGeneratedFromClaudeFactory
invalidusernam3 8 hours ago [-]
Off the top of my head I would have thought zero width characters (eg: U+200B, U+200C) making some unique identifier sprinkled in amongst the output. But obviously far from foolproof since they could simply be removed.
gajus 21 hours ago [-]
Most likely watermark will be proportional to the input/output ratio, i.e. if you input a long document and ask to make edits, it will not attempt to watermark it. On the other hand, if you provide a tweet and ask it to write an article, that will include watermark. Just a guess (and yes, it feels flawed)
kkukshtel 5 hours ago [-]
I think a lot of the examples below are projecting more complicated options, but it could also be something just as simple as using a word with a hyphen in it every prime-numbered sentence. Or any other "puzzle-y" pattern.
nprateem 13 hours ago [-]
Load-bearing==claude
siva7 21 hours ago [-]
I can tell you how: Claude produces a huge wall of text with jargon ridden bullshit and invented terms no human subject matter expert would seriously use and overuse.
mchusma 6 hours ago [-]
Many good comments here. It’s somewhat common for me to voice record say a blog post of product updates, more like a ramble. Then have Claude clean it up. Then argue back and forth about certain things until it’s good, then make a final pass sometimes to change a few key words. This is incredibly different than pure ai text. Presumably it will show as ai generated here, even though I would argue it is not really. So I can’t use Claude for this usecase anymore.
I think the solution is assume everything is ai generated unless told otherwise and rely on authorship/brand as a sign of quality.
gwillen 3 hours ago [-]
The watermark is contained in the choice of output tokens. If you exert so much editorial control that Claude has no meaningful freedom in choosing tokens, then it's going to fail to watermark the text, unless the text is _very_ long (in which case even a very low-bit-rate watermark will eventually accumulate enough bits to be positive.)
MattSayar 3 hours ago [-]
I've done this before at work, and I feel true ownership of the output after this workflow. Moreso than when someone from a marketing team publishes a blog with the CEO's name as the author.
margalabargala 4 hours ago [-]
You can do something like "here's my website. Read it, then try your best to write in my voice. Do your best to avoid common and uncommon tell-tale AI-isms"
If you're putting the work you say in, the result won't be obviously distinguishable. Obviously, from some of the things that get posted here, that last sentence is too much for most people to bother adding to their prompt.
no_multitudes 3 hours ago [-]
Simply write your own posts if you don't want people to think they are AI generated.
sfink 4 hours ago [-]
Perhaps. But it seems like your beef is not with the presence of watermarking, it's with what people will use that watermarking for. You're not directly harmed by that blog post being labeled as AI generated. In a hypothetical (but unfortunately likely) world where everything has passed through an AI's digestive system, nobody would care.
In the meantime, it is true that this takes something away from you. But it's something you were only recently given. Now you're not given quite as much, but readers are given a little more (or rather, there's less being taken from us!)
> This is incredibly different than pure ai text.
Ok. But it's still incredibly different from pure human text. I guess the question is which provides more value? Providing the information "this text is AI watermarked" to readers? Or allowing creators to lie and claim that AI processed text was 100% human generated? I agree that people assuming that "has AI watermark" == "is AI slop" is incorrect and causes some amount of harm, but having the watermarks also pushes back on a large amount of harm already being done.
(Personally, I'm skeptical that these watermarks will ever hold up to adversarial attacks, and they haven't claimed that they will. So I think it's the usual "casual liars will be caught, determined liars will get an additional thin veneer of respectability".)
altmanaltman 3 hours ago [-]
I would push back on the fact that nobody would care. It is clear that platforms are increasing creating AI-generated as a category and it will only get more precise over time. Platforms that have built trust/aura over what type of content they host will resort to this more as they get more flooded with ai generated low effort content. I recently wrote a blog on this actually: https://decodingvibes.com/blog/aura-and-the-backlash-against...
altmanaltman 3 hours ago [-]
But it is not incredibley different than pure ai text tho, it is literally the same as pure ai text.
I understand you're saying since you "worked with it", it is not ai generated but if you still use the final output verbatim, the writing itself is LLM generated purely.
You want to share the output by it but also position it as not ai output. But that's fundamentally dishonest.
Furthermore, if you think your approach actually creates value and can be judged on its merit, why not disclose its ai written? If you think that will make people think your content is bad then you should see that as feedback and maybe not use AI since readers don't like it.
benrow 21 hours ago [-]
I've heard that this kind of watermarking process works by biassing the statistical sampling towards a partition of the set of possible next tokens (red set and green set), at each position. It might only be a slight nudge each time, but over a sequence of tokens, the likelihood of repeating the bias by chance is increasingly improbable.
The bias is different for each position and follows a defined RNG, seeded somehow predictably.
Can be either an open algorithm, or not. If not open, then an API could be provided to determine if text is watermarked or not.
How it applies to code - maybe it could be a subtle nudge to symbol names, etc, I'm just speculating (I only read about this in passing very recently).
m-chrzan 21 hours ago [-]
There's a computerphile video (https://www.youtube.com/watch?v=XZJc1p6RE78) with Dr. Mark Pound explaining a paper by John Kirchenbauer, Jonas Geiping et al. (https://arxiv.org/abs/2301.10226) that described a method for watermarking LLM output like this. It's not directly stated anywhere in the Claude support article that this is what they're using, but the properties of the watermark described seem to point to this method.
londons_explore 9 hours ago [-]
The bias has to be small enough that if you ask an LLM to repeat some passage of text like the national anthem, either from the training data or from the prompt it doesn't change random words.
Gotta be hard to tune that.
8 hours ago [-]
metalcrow 18 hours ago [-]
Based on my understanding, it can only be applied to code in very limited ways: docstrings, variable names, string literals. The code itself can't really have tokens changed to another equally correct token (the foundation of the watermark) because then the code breaks! And the few places that you can do so are likely erased by formatters anyway.
matherial 4 hours ago [-]
You have less latitude than in prose, but I think you're underestimating how much can be changed without causing breakage. For example, the most likely sequence might be "if (!foo) bar; else baz;", but you can also say "if (foo) baz; else bar"; substitute "foo == 0" / "foo != 0" for even more variety. Similarly, "foo = 1; <NL> bar = 2;" can be output as "bar = 2; <NL> foo = 1;".
Keep in mind that the LLM "sees" the previous (tweaked) output and picks what makes sense based on that. There are few situations where a perturbation like that would be unrecoverable, and I assume these situations also correspond to a huge probability difference between the most likely completion and the second most likely one - in which case, the watermarking algorithm can choose not to touch the token.
cassianoleal 21 hours ago [-]
> a defined RNG, seeded somehow predictably
So, an NG?
olmo23 7 hours ago [-]
PRNG
IshKebab 20 hours ago [-]
If it's based on position mod 2, wouldn't inserting or deleting (or splitting/merging) words every now and then trivially defeat it?
If it is based on position mod 2 then wouldn't inserting/deleting (or splitting and merging) words every now and then defeat it?
mhjkl 7 hours ago [-]
All big LLMs already visibly watermark all their text with easy to detect annoying phrases and turns of speech that everyone is already sick of hearing. Why do AI companies keep making their products worse to appease anti-AI, it’s not like they’ll suddenly start supporting it if you do so. If you’re worried about European customers, just relax your firewalls to let more VPNs through, if the productivity boost is high they will use it anyway if their rules keep crippling their own models
IsTom 7 hours ago [-]
> if the productivity boost is high they will use it anyway if their rules keep crippling their own models
Individuals maybe, companies won't and that's where most of money is at.
mhjkl 7 hours ago [-]
This will make more money for Claude from individuals unofficially acting as meat puppets in more inefficient workflows
WD-42 5 hours ago [-]
Is it to appease anti ai or is it a method they will use to avoid training on their own slop?
KronisLV 24 minutes ago [-]
> Claude models launched in the EU
Could we get an Anthropic subscription for Claude Code with data residency in the EU, so we don't get robbed blind by AWS Bedrock et al., but can have a monthly subscription like with the regular US option?
jonplackett 12 hours ago [-]
We need to just stop pretending we can reliably tell if plain text is written by an LLM.
It’s just not a reasonable ask.
nunez 6 hours ago [-]
Now that the EU mandated watermarking, the point is that services (or browser extension developers) can add their own detectors to make AI-generated text obvious. It won't fix AI in print, but most of the problem is online anyway.
jonplackett 3 hours ago [-]
These things are trivial to remove though. And the whole point of it is they will also make the _detector_ available so you can then also check if you successfully removed it.
It’s not a solvable problem.
JohnKemeny 10 hours ago [-]
True, but what you can do is a one-sided guarantee. If it bears the mark, it is likely generated (or someone deliberately made it look generated).
Thus, if a news article, research article, book, student paper submission, blog post, HN comment, etc, bears the mark, it could be automatically flagged as such.
It helps detect low effort slop.
---
Caveat. If you write your own creative work and send it to Claude for "cleaning up grammar", it might insert the watermark.
jonplackett 8 hours ago [-]
The problem with pretending is that people who k ow what they’re doing get away with it while people who don’t (and don’t even use ai) get unfairly accused of using it.
There just isn’t enough information in plain text to do this and we should stop pretending there is.
If we need to verify something isn’t made with ai then we need other ways of doing so - eg looking at a document edit history, doing it as an exam, oral defense.
There are options! But pretending you can tell if text is ai will only catch out people who make no effort to hide it and will inevitably have false positives.
DanielHB 9 hours ago [-]
It seems like it would be so low effort to bypass, especially when you can just train a system (maybe even another LLM) using the watermarker validation from Anthropic themselves.
Might help with students and low-technical people passing off work as their own, but any industrial scale slop-generator should be able to bypass it trivially.
Art9681 7 hours ago [-]
Not might, will. Whether enough text is present or not to go over the detection threshold is in doubt. But the "score" will never be zero, even for human written text.
someguynamedq 6 hours ago [-]
will insert a watermark
asnelt 9 hours ago [-]
> If it bears the mark, it is likely generated (or someone deliberately made it look generated).
One could even say, the mark is load-bearing.
akersten 20 hours ago [-]
So my code that Claude makes, which previously was using the best (most probable) tokens for the job, will now be getting worse in random positions, to appease a voluntary EU suggestion. Love that.
neuroticnews25 11 hours ago [-]
They aren't using greedy decoding, there's enough randomness in sampling to swap some with independent signal.
akersten 7 hours ago [-]
Purely greedy or not, there is some measure of "goal outcome" that was previously being solved for with the token selection function, and the goal was "complete this text with the best (surely, otherwise what are we doing?) next part, and sometimes the best next part is a little bit random just to keep things interesting"
Now the goal is either "identify the meaningless interesting bits and swap them out with 0% loss in the direction of the original goal," or "perturb some small selection of the output towards my secondary secret goal of watermarking the text."
It would be quite impressive if they managed to identify with 100% accuracy the tokens that "don't matter" and are free to swap with whatever signalling tokens encode the AI scarlet letter, but most likely they are not 100% accurate, and that means the output is worse off than without the watermarking logic.
neuroticnews25 3 hours ago [-]
What you're saying sounds intuitively true and from what I've found modern watermarking methods measurably rise perplexity by 1-3% [0]. Gemini convinces me it doesn't matter and doesn't compound over long contexts though. I would love to see HN experts opinion.
Oh, how I laughed. That was never your code, my friend.
akersten 7 hours ago [-]
Please don't let the arbitrary selection of phrase distract you from the substance of my argument: a product that I pay for is at best no better due to this change, and highly probably worse. Why am I paying for a tool that is beholden to clandestinely satisfy some far away master?
mikro2nd 6 hours ago [-]
Good question! Why are you paying for some tool that has always been beholden to some faraway master's opaque agenda?
someguynamedq 6 hours ago [-]
We live in a world where outcomes are often more important than process
Banditoz 2 hours ago [-]
I don't know, why are you? Who's to say Anthropic wasn't already modifying output in some way for optimization or some other reason? These models are incredibly opaque.
lelanthran 5 hours ago [-]
> Why am I paying for a tool that is beholden to clandestinely satisfy some far away master?
Why were you doing that before watermarking?
Same answer.
Pavilion2095 4 hours ago [-]
I assure you, this is one of the mildest things they do after training before the model reaches you.
fwlr 6 hours ago [-]
Because the tool is made by a corporation that is subject to regulation by a government, and that government has decided it’s in the best interests of society that the tool be limited in this way.
HatchedLake721 7 hours ago [-]
Who's is it?
mikro2nd 6 hours ago [-]
That's just the point: nobody knows! It might be my code. It might be your code. Nobody knows who they stole it from.
MagicMoonlight 11 hours ago [-]
[dead]
KETpXDDzR 23 minutes ago [-]
I wonder how the caveman skill will affect this. Also, you can probably put the output of Claude in another LLM to get rid of the watermark.
bramhaag 8 hours ago [-]
I cannot wait for the inevitable "I've always used Claude watermarks in my writing, even before we had LLMs!" when someone gets caught using an LLM.
andai 3 hours ago [-]
If I understand correctly, this means that any text with the "watermark" is legally uncopyrightable, including code.
Not an copyright attorney, but color printers have watermarks. That's never been an obstacle.
luxuryballs 2 hours ago [-]
what if the output is downstream from copyrightable work? wouldn't the LLM touching it wash that off if this was the metric used?
mcv 2 hours ago [-]
I think it's definitely important that any AI generated content can be easily identified as such. I think the new EU law that requires that, has too many unnecessary exceptions.
So great that Anthropic is doing something about this, but it's not clear what their watermark exactly is. How do I, as a user running into some content online, know that it's generated by Claude? What is their watermark?
It sounds to me like they create the pattern in the regular text of the content, which sounds interesting, but also odd, unreliable, and may limit the content you can get out of Claude. Will it subtle change the words in order to hide this pattern in it? I don't know if that's something anyone wants.
allthetime 1 hours ago [-]
"Will it subtle change the words in order to hide this pattern in it?"
I'm assuming that's exactly how it works. How else could it?
aabhay 22 hours ago [-]
I have had a hunch for a while now that (in addition to these tools), Anthropic has actually leaned in to Claude's distinctive manner of writing since it makes the text more obviously AI generated and thus less susceptible to misuse.
That's not necessarily the same thing as a markov-style fingerprint but it could be a correlated factor.
pjm331 21 hours ago [-]
I had a similar thought but I assumed they leaned in because it improved performance on coding or something like that
LoganDark 21 hours ago [-]
I suspect it's because of alignment concerns. The more deeply they can integrate their principles, the harder it'll be to misuse. Or at least that's the idea.
kingstnap 21 hours ago [-]
It could also partly be a byproduct of examples of claude writing being in the dataset, which of course anthropic has lots and lots of and they do train on.
r_lee 11 hours ago [-]
no way. there's just no good excuse for why "load-bearing" and "worth flagging" are everywhere now, I've pretty much never seen that in the wild before
noman-land 19 hours ago [-]
It's pretty trivial to command it to not speak that way. That's one of the first things you should write into the prompt. What style you want it to write in. Make it use a very concise and dry academic style with no overt LLMisms, melodramatic or flowery language, or metacommentary.
breezybottom 7 hours ago [-]
People have been posting some variant of this comment for three years, and it's no more true today. Ever notice that the "prompt engineer" career hasn't materialized?
someguynamedq 6 hours ago [-]
Prompt engineer is a requirement within every serious job now, not a job in itself
noman-land 6 hours ago [-]
Regular engineers still exist.
fwlr 6 hours ago [-]
A surprising number of people are worried that the code they don’t read will be imperceptibly different.
simonw 21 hours ago [-]
An interesting factor of this is competition.
If Claude was the only model family they could ship a change like this and users who want to cheat (or don't like watermarks for other reasons) would just have to put up with it.
In a world with many different competing models, the risk of losing customers to other providers over this is much more real.
Maybe they've looked at the numbers and the portion of people who clearly use Claude to cheat on examples etc is so tiny that losing them to other providers isn't a problem?
lhd1 14 hours ago [-]
Scott Aaronson spoke about this in a colloquium where he said that this was mooted at OpenAI before the decision was made by Altman to not implement it for the reasons you describe.
I’m more worried that this will degrade performance. I want the best results from a model, not the results that fit a constraint that’s not defined by me. Any increased cost or latency is also unacceptable.
I guess whoever is the policy maker is assuming that some protection is better than none and that most people will not reach for such tools.
nprateem 13 hours ago [-]
Either that or they want to comply with the EU AI Act when it affects them.
Dkuku 1 hours ago [-]
Thats why recently it pushes so much comments in code - it has to squeeze the watermarks somewhere
Groxx 2 hours ago [-]
>Content generated by Claude may not carry a detectable mark if, for example: ...
>A file’s metadata was stripped through format conversion, re-saving, screenshots, or other means
Ah. So what essentially every single consumer-oriented media host does. Gotcha.
I fully recognise this is a hard problem, but hopefully metadata isn't the only method for media. Standard procedure is to shrink files for storage and privacy reasons, and non-visual metadata goes out the window by default.
ethin 21 hours ago [-]
Can someone help me understand how exactly this watermarking of text works?
Given that text is, well, text, and not some kind of binary format, I don't see how any watermarking can work unless you insert characters which are invalid under Unicode. I further don't really understand how this won't be perceivable by assistive technology (the "watermark" will just appear as either unreadable characters, or if the watermark is mixed thoroughly enough into the text, it will scramble the text to any speech synthesizer and will make it really really obvious). Thus, I don't see how this wouldn't be insanely trivial to remove. And this is before we get into things being put on the clipboard. Sure, I can press the "Copy" button at the end of each response, but what I can also do is manually select the response and copy it, or only copy partial selections, or any number of other things. How does this "watermark" (or any "watermark" technology) take into account this?
So, really, to summarize this: I see no way of this actually being technologically achievable unless we revise the very core of how computers work and encodings for textual information. So I'm very curious as to how this is actually supposed to work.
benrow 21 hours ago [-]
Have a look around for token biasing, or green lists. It's based on a nudge to the choice of the next token (which can always be drawn from a set of possibilities which are all probable enough).
At first I thought this approach was just the "LLM flavour" of writing, but it's way more subtle, especially as the bias is applied uniquely for each token position.
ethin 20 hours ago [-]
Yeah, will do, this sounds interesting since I'm not entirely sure how this would actually be reliable to any degree. Thanks for the help, not sure why I got downvoted since I was genuinely curious.
fl0id 4 hours ago [-]
there have been papers about it, it works
resonantjacket5 21 hours ago [-]
it's a statistical way. like for example maybe in your above paragraph claude maybe writes "Thus, I don't see how this wouldn't be insanely <easy>(instead of trivial) to remove" and then also says like "And this is before we <analyze> things being put on the clipboard." or maybe the i just says the word "the" in a certain pattern or frequency.
you can then consistently like figure out if it was claude that wrote the sentence. it is easy as you noted if you just get another ai to read it and then rewrite it.
MagicMoonlight 11 hours ago [-]
[dead]
plutokras 55 minutes ago [-]
This feels like obvious setup for regulatory capture. It won't be long before missing "safety" watermarks are cited as the pretext for restricting Chinese models.
taormina 2 hours ago [-]
New models will mark AI-generated content from day one. Claude models launched in the EU on or after August 2, 2026 will support machine-readable marking at launch. Generated text will carry embedded watermarks, and generated files will include digitally signed provenance metadata where supported.
So, they’ve been doing this for over a week without telling anyone?
edg5000 4 hours ago [-]
This is outrageous. I hope only Anthropic will do this. Are they going to disclose at least the specific Unicode whitespace characters used for the watermark? Or will they use some other trick?
If I heavily edit LLM output, will this still hold the watermark?
You really can't make this stuff up, it doesn't make any sense.
secretsilver 4 hours ago [-]
I don't think they are using any invisible Unicode or metadata stuff.
What happens is that AI selects similar words based on a random process.
Something like "The company had a large/big/substantial advantage".
It chooses between these words, and over a longer piece of text, the pattern will start showing, like a "choice A → choice C → choice C → choice B → choice A".
The normal-looking text will actually be a fingerprint living in the form of statistics.
I think Claude will be sharing these patterns to third parties for AI detection.
wavewrangler 34 minutes ago [-]
Anthropic, do ye come load-bearing gifts?!
graypegg 5 hours ago [-]
Huh... I wonder if some big version of a bloom filter would work as well. Hash all output text, probably in chunks of a couple tens of tokens each (that would need tweaking to find the most useful hash input length I guess), and smash 'em into a bloom filter. Every month or something, Anthropic releases a giant file containing whatever huge length of bytevomit would have to be used to get an acceptable false positive ratio. (In terms of bloom filters! Meaning: still far from a perfect ratio.) Maybe one for each model they provide or something?
Then at least you could have two weak-postive signals, and a strong-negative signal. (Though one that only fits precise chunks of tokens) I'm sure I'm missing something here, but my groggy morning brain thinks that doesn't seem too bad.
wowokruyi 1 hours ago [-]
If I have Claude directly translate my original words, does that mean my original words also get watermarked?
0x_rs 1 hours ago [-]
It's safe to assume so. It says "the output can carry a Claude mark even if the underlying ideas, text, or data originated from another source" specifically about translations among other tasks. It's not strictly about generation but processing, and what level of processing is involved is entirely arbitrary, that is to say: you must feed your text into proprietary black box machines to figure it out, because it's not something you should be supposed to tell otherwise.
drra 41 minutes ago [-]
yes, it's likely quite mechanical token distribution pattern so your text is going to be watermarked.
drnick1 21 hours ago [-]
Seems like an awful idea. I hope that that "watermark" will soon be discovered, reverse-engineered, and that tools to remove it will appear.
cassianoleal 21 hours ago [-]
I hope all models adopt it.
DaSHacka 11 hours ago [-]
Thankfully, there are a variety of Chinese models that never will. I think we all know that in a few years, they will also be the only relevant offerings on the market, due to not being bogged down with over-zealous ""safety"" footguns.
wpietri 9 hours ago [-]
Your theory is that the Chinese government is thoroughly uninterested in safety or prosocial controls?
dannyw 7 hours ago [-]
Amongst Chinese labs and netizens, there's MUCH less belief/mindshare on "AGI = existential risk to humanity", "paperclip maximiser", and similar lines of thinking. AI is seen more as just a technology, and less like a scary boogyman.
Whether that's right or wrong, I'll leave to you, but there's huge differences in perspectives, and if you only get your news from Western sources and communities (and companies), you're in a bubble too. A different bubble, and arguably a more porous one, but still a bubble.
wpietri 6 hours ago [-]
The paperclip boogeyman is not the only reason, and probably not the biggest one, that models get safety/content constraints, however useful it is as a PR distraction. I agree Chinese models are different right now, but I think that's a function of their novelty and desire to compete globally.
For a taste of where I think things are headed, try asking Chinese models about Tiananmen [1]. And then take a look at the Chinese government's approach to pretty much anything that they think reduces security or social harmony. I find it hard to believe their models will be the one exception to that over the long term.
Other models will end up diffusing it and making the signal indeterministic and irrelevant.
j16sdiz 2 hours ago [-]
The true reason is legal.
> Anthropic has signed the EU AI Act's Article 50(2) Code of Practice on Transparency of AI-Generated Content, ...
izonu 21 hours ago [-]
> We’re also working to enable users and other third parties to detect Claude’s embedded watermarks and provenance metadata.
This seems to be similar in execution to Google's SynthID. I hope they release actual code the technically proficient can use, unlike SynthID which can only (afaik) be queried with Gemini's UI.
gajus 21 hours ago [-]
The moment Google announced SynthID, the first domain I bought was deSynthID.com
Several open-source projects have already proven SynthID to be ineffective.
moebrowne 12 hours ago [-]
There are many free lock picking tutorials, but yet locks are still effective.
mjuarez 4 hours ago [-]
Like others have said, it's not reasonable to ask this.
I propose we defeat this with the obvious: Simply, figure out what are some of the markers Claude and others will use for these tools, and sprinkle them randomly on everything we type or produce, all the time, 100%. If users flood the tools, and everything returns as AI-generated, then the tools become useless.
jobigoud 2 hours ago [-]
The markers aren't words that you can sprinkle in your prose, they are a statistical bias in the selection of the words.
FeteCommuniste 4 hours ago [-]
Why would you want AI content to be indistinguishible from human output?
0x_rs 6 hours ago [-]
Is the detection mechanism going to be open, free, and possible to run locally without prostrating to an opaque third-party company that will do whatever they want with the text content provided (including using it for training), and take no responsibility in case of false-positives for which there can exist no proof or evidence against by the victim? This is another useless, if not actively harmful, performative EU regulation, for which they ought to take the full blame despite the fact "AI" companies have been researching and working on watermarks, including in text, for a while now. Just copy-and-paste everything you see and let a machine decide for you if what you're reading is slop or not. Real propaganda machine doesn't care about inane rules and won't waste time with gimped mainstream models either.
>Claude may not be the original author. People often use Claude to proofread, translate, summarize, or convert files. The output can carry a Claude mark even if the underlying ideas, text, or data originated from another source;
Such models already struggle not making any unnecessary or unwanted changes to a corpus, this makes them unable to by design.
letmevoteplease 5 hours ago [-]
It is of course a stupid regulation, but the upside is that it will probably accelerate growth in usage of open models that are not adversarial towards the user.
edg5000 4 hours ago [-]
I like your optimistic way of thinking. There is a lot of truth to it.
Godsend69 4 hours ago [-]
[dead]
Sha1rholder 2 hours ago [-]
Who would take the responsibility of misjudgements then?
hoppp 4 hours ago [-]
How will they watermark code? I can understand encoding something in free form text but for example I ask it to generate a react component, will it embed watermark in the typescript code?
bargainbin 3 hours ago [-]
console.print(“these logs are a load-bearing seam, do not remove”)
It’s foolproof, I tells ya.
Schlagbohrer 5 hours ago [-]
I've long thought we would have some sort of verified-point-of-origin for data using a hash or cryptographic seal of some kind. I don't know the precise technical language for that but some metadata traveler that can verify the data has not been edited after creation.
someguynamedq 6 hours ago [-]
Notice the subtle capture in you having to use Anthropic to identify Anthropic's watermarks? Smart of them
matherial 4 hours ago [-]
I guess this is where our true colors show. There's a significant contingent of HNers who always dunk on LLM text detectors and claim that they can't possibly work, that they ruin careers, etc. But now that a lab says "OK, we'll add a real watermark", the reactions are overwhelmingly that it's still somehow wrong.
Why do feel so entitled to being able to pass LLM-generated text as our own? I get that a lot of techies aren't good at writing. I also see a lot of tech hustlers who like to use LLMs to fake human connection and compassion - I've gotten LLM-generated recruiting emails that talked at length about how the recruiter "valued" my work. Just because we found a "cheat" button doesn't mean it's wrong for others to want to know.
Yes, LLMs are great. So is transparency. If you think an LLM writing is your new superpower, wear that badge with pride. It might mean you will lose some business from LLM haters and win some other business from like-minded customers. C'est la vie.
pickleRick243 2 hours ago [-]
"A detected mark provides a signal that content was processed by Claude, but is not fully conclusive."
"Lack of a detected mark doesn’t mean the content wasn’t AI-generated or processed."
godd2 4 hours ago [-]
> Just because we found a "cheat" button doesn't mean it's wrong for others to want to know.
One difference perhaps is that you think using LLMs is cheating, while others do not.
davisr 2 hours ago [-]
Having a ghostwriter in your ordinary life is absolutely cheating. It's letting someone, or something, else write for you and pass it off as your own. Personally, I'm insulted any time someone sends me LLM-generated media.
FeteCommuniste 4 hours ago [-]
It saves the non-anglophones the bother of learning to write readable English, so there's that.
fl0id 4 hours ago [-]
they might be two different sets of people
swedishuser 6 hours ago [-]
I wonder after how much editing an LLM output wont be reliably detectable? And what the EU law even says about this. I find that a good LLM workflow can be to generate outlines that are then edited pretty heavily manually to fit into whatever context it will be published in.
jgilias 5 hours ago [-]
The more they fiddle with the autocomplete system, the more they move away from the autocomplete faithfully producing the completion I need. The more it makes sense to move to an open weights model not served by them.
edg5000 4 hours ago [-]
The pull is strong, but it will take a few generations of hardware before 1 TB becomes attainable without needing 100k, but more like 10l (kinda like the value of a car, defendable as a job expense).
Anthropic has this strong repulsive effect in the way they operate, I wonder if they'll be around for long, it's hard to say at this time. The competion is fierce, so there isn't much room for shenanigans at this stage.
nojs 8 hours ago [-]
No mention of what data they are specifically encoding. Will it be like printing dots, traceable to the exact account that generated the text?
edg5000 4 hours ago [-]
From what I read in the comments, the model will be more biased towards certain words that otherwise would would have a very simmilar chance of appearing (e.g. very simmilar words that would not really alter the meaning of the text). If this actually works the way I think, that's a really sneaky way to hide data in data. It's clever, but hostile towards the user.
jrhey 6 hours ago [-]
I wrote about how this might work here without invisible characters:
So we will use open source model to copy text from high end models ?
darrinm 4 hours ago [-]
Are they doing this as a means to detect distilling? What kind of encoding would survive that process?
raincole 7 hours ago [-]
Why would I want a stochastic parrot that intentionally speaks less probable tokens?
Anyway, I've found a magic line that can be copied & pasted to the comment sections of most OpenAI/Anthropic news threads. This one is no difference.
The magic line:
> Doesn't matter; have DeepSeek.
normalaccess 5 hours ago [-]
Would this impact distillation? Reminds me of the fake roads map makers would add to their maps to detect copying.
andreypk 8 hours ago [-]
Interesting technology. I wonder what else this could be used for beyond AI-content detection — e.g. provenance, model attribution, or tracking how generated content evolves through edits and transformations.
lorenzohess 22 hours ago [-]
> Generated text will carry embedded watermarks, and generated files will include digitally signed provenance metadata where supported.
This should make it easier to catch cheaters who use Claude, right? Unless everyone runs their artifacts through some watermark and metadata sanitizer?
uncivilized 22 hours ago [-]
As long as they’re in the EU.
nonestdeus 21 hours ago [-]
From the linked article
> Regions. Marking will apply to output from supported models wherever Claude is offered, worldwide.
bossyTeacher 22 hours ago [-]
> Unless everyone runs their artifacts through some watermark and metadata sanitizer?
It will happen if Claude tampers the text. Guaranteed.
pixl97 20 hours ago [-]
Text is too low bandwidth to classify reliably without lots of false positives. Especially as people start talking like LLMs.
reasonableklout 12 hours ago [-]
The approach Pangram has taken which works pretty well is to simply lower the recall a lot but ensure the precision is very high. Which means potentially high false negative rate but low false positive rate.
coolboydev 4 hours ago [-]
This shows just how important open-source language models are.
case540 22 hours ago [-]
I don’t like the idea of hacking a response to contain a watermark. I also don’t like the idea of false positives detections coming directly from Anthropic. If people read more AI generated content, people will probably start writing more in that style
stranded22 13 hours ago [-]
The amount of times ‘delve’ appeared in general conversation in the last couple of years shows the influence LLMs have on society.
pixl97 20 hours ago [-]
I have no idea why you were down voted for this. Language is alive and people adopt it from sources they hear a lot.
ack_complete 18 hours ago [-]
Moreover, what if you quote text that happens to have been generated by Claude, does that bump up the AI-ness score of your source file or document?
reasonableklout 12 hours ago [-]
The flip side of this is that if AI-generated content becomes reliably identifiable and carries a stigma, then people might deliberately change their styles to be more diverse and human.
One example I've seen are junior employees at my company deliberately adopting a lowercase/less punctuation writing style so as to stand apart from AI.
mateocafe 8 hours ago [-]
Seems to me like this creates a huge incentive to game the watermark. Also, how does it prevent having AI generate the text, then the user copy-paste it into a clean document?
nsvd2 8 hours ago [-]
The watermark is in the text. If you copy the text you're copying the watermark which is part of the text.
Computerphile on YT has a video explaining how models can fingerprint the text they produce. Essentially they modify the probabilities of word choice slightly in a predictable way.
mateocafe 7 hours ago [-]
Interesting, thanks for the explanation, will check out the video.
This would imply that a positive watermark signal is likely (but not guaranteed) to be AI generated. Also implies that a negative watermark signal is not necessarily void of AI generated text. This would create a problem if people start to trust the watermark as a heuristic, as the ability to critically evaluate the text is replaced by the search for a watermark.
Seems to me all of this is really trying to solve for "is this text bullshit" or not, which would require a different solution.
partsch 7 hours ago [-]
Perhaps one should start by looking into how the providers of LLMs obtained the training data.
luciana1u 6 hours ago [-]
the watermark is the least interesting part. the interesting part is that we're now arguing about whether a machine's handwriting is legible enough to count as a signature.
someguynamedq 6 hours ago [-]
Watermarking text is impossible and a fool's errand
If we go by "fool's errand" as "needless or profitless endeavor", https://arxiv.org/abs/2303.11156 may already be a good enough answer to the paper you cited, so their work is already laid out for them. The green token idea is thoroughly attacked with much more effective techniques than those in the original paper through "recursive paraphrasing". Among some hypotheses in the paper, one is particularly interesting:
>These experiments provide empirical evidence that more advanced LLMs can lead to smaller TV distances. Thus, based on Theorem 1, reliable AI text detection would become increasingly difficult
dalemhurley 21 hours ago [-]
People with dyslexia and dystrophia, commonly use LLMs to proofread content. Even Anthropic admits this is a limitation.
stranded22 13 hours ago [-]
Yes, I’m audhd and dyslexic.
I am cancelling my Claude max 5x subscription and moving to ChatGPT pro. I have difficulty enough trying to ensure my meaning comes through correctly, along with everything else; to now have to look out for/analyse watermarks too?
I feel shamed enough by society, thanks Anthropic.
Terretta 8 hours ago [-]
Paradoxically, one of those two firms puts considerably more effort into accommodating such differences, and the other has signed the same EU law and just hasn't performed as well rolling it out.
Both points suggest your subscription support was well chosen before.
If it's only proofreading text you've written, the changes will be minimal enough that watermarking seems impossible to me.
breezybottom 7 hours ago [-]
How could "your meaning" come through if a computer is writing it?
bramhaag 11 hours ago [-]
How exactly does this impact proofreading? You can manually apply the suggestions (typo here, unnatural sounding sentence there, etc.) the LLM gives you to your own content, and it would stay watermark-free.
Unless with "proofreading" you actually mean having the LLM write your content for you.
matheusmoreira 14 hours ago [-]
People with executive dysfunction too. LLMs bring execution costs down to near zero and are therefore assistive technology.
Dilettante_ 11 hours ago [-]
"This is my emotional support gun. It makes me feel safe despite my CPTSD and is therefore assistive technology."
michaelolenick 4 hours ago [-]
It's inevitable there will be false positives, inevitable they'll do reputational or economic damage, and inevitable plaintiff attorneys will sue on the behalf of people damaged. Making it worse, it's product liability blended with defamation. Anthropic should've told the EU to pound sand and geofenced off Claude. If they don't want to live in the dark ages, elect smarter people.
padolsey 9 hours ago [-]
Is this just to appease regulators? They surely know this won't work in the long run.
tiahura 6 hours ago [-]
You mean like with printer watermarks?
aniceperson 8 hours ago [-]
That's not only load bearing — it sustains the need to detect AI content
hparadiz 21 hours ago [-]
I know you're gonna read this so I'll be blunt. This is bad for your brand.
LEDThereBeLight 19 hours ago [-]
Reactionary emotional advice does no good, it just makes people want to hold their positions more defensively. If you care enough to say something, care enough to say it with reasons that might shift someone’s perspective.
SubiculumCode 2 hours ago [-]
Go ahead and follow your own advice. You made a claim. Now back it up.
Laurel1234 5 hours ago [-]
It's not up to Anthropic.
charlieyu1 5 hours ago [-]
All for more surveillance.
partiallypro 3 hours ago [-]
My question is what's to stop Google or any competitor from using watermarks of Claude or OpenAI from degrading the rankings of sites that use it but ignore or even reward sites that use Gemini. Seems like an easy thing to do for competitors, and maybe an unforeseen side effect of these types of things or regulations.
lukewarm707 21 hours ago [-]
no thanks.
jp0001 21 hours ago [-]
OpenAI has been watermarking their images with C2PA for some time.
cybice 4 hours ago [-]
на ху я?
fidotron 7 hours ago [-]
The SV obsession with neo Kabbalistic nonsense will get a whole new burst of energy from this.
someguynamedq 6 hours ago [-]
Please say more
GrayHerring 21 hours ago [-]
Wasn't enough to play cat and mouse with ad removal, now we can also do the same with watermarking.
jp0001 21 hours ago [-]
You could flip bits in the font itself, but I'm really wondering how portable this is.
dejanseo 11 hours ago [-]
> "Claude models launched on or after August 2, 2026 support marking at launch."
No Anthropic model has been launched in August.
chasing 5 hours ago [-]
Won't there instantly be tools to detect and remove/obfuscate these kinds of watermarks?
mucha 1 hours ago [-]
Absolutely! You can avoid the watermarks by writing or rewriting everything yourself.
chasing 1 hours ago [-]
I'll just have my AI do it...
svaha1728 21 hours ago [-]
I expect a “Prettier” for AI generated text in the near future.
morkalork 7 hours ago [-]
If they can use a cryptographic key to sign the text, will they generated keys unique to users? Seems like they'd be able to trace back to which accounts were used to say generate scam dialogue, threats, propaganda or bot content on the open internet?
hbn 19 hours ago [-]
If the western AI companies are forced to comply with this type of BS, and develop their models to do their job while balancing a book on their head and hopping on one foot, the Chinese models just got a free pass to completely dominate the frontier.
EU regulation does it again!
Mashimo 11 hours ago [-]
If the Chinese want to sell to EU customers, they probably have to do the same.
mhjkl 7 hours ago [-]
I’ve seen Chinese open weights models say “can’t use this if you’re in Europe” in their licenses, so I doubt they would invest too much into complying
Mashimo 7 hours ago [-]
You mean the model itself? Yeah, probably not needed if you self host. If they run a service that they sell and they are a bigger player like Alibaba, I doubt they can just ignore it.
DimitriBouriez 11 hours ago [-]
What's the problem, really? Given the direction the U.S. has been heading in recent years, I wonder what really sets it apart from China. Europe needs to maintain an equal distance from both the U.S. and China.
krzyk 5 hours ago [-]
How is that BS?
I would prefer to know if given content was generated with LLM. This is information, and information should be free.
breezybottom 7 hours ago [-]
Less AI slop sounds like a win to me. Let China drown in it.
VCFundedGenYer 7 hours ago [-]
> When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won’t see it, and it doesn’t change the meaning, quality, or readability of Claude’s response.
I feel like this is FUD. If you copy text from Claude, Ctrl Shift V it into VS Code, the IDE will give up the ghost on if weird characters are in there. And it's not like Google suddenly invented new letters or fonts either.
Practically speaking, I feel this is Google publishing misinformation.
amelius 22 hours ago [-]
They should just replace the spaces by one of Unicode special space characters.
Can it be circumvented? Of course. Will most people go through the trouble to circumvent it? No.
A_D_E_P_T 21 hours ago [-]
If it's that simple and obvious, you'll have 10 "Remove Claude Watermark" web-apps by the end of Day 1. Most of them coded by Claude.
Hell, it'll probably happen no matter how sophisticated their watermark is. There's no watermark in text that can't be detected and removed, and no text that can't be converted to generic keyboard ASCII.
amelius 21 hours ago [-]
You forgot about the cases where (1) people don't care, (2) people want to say "I used an LLM for this". I'm convinced that those cases happen more often than you think. Why not cover them with a simple mechanism? It's also in the interest of AI companies who don't want to train on AI output.
pixl97 20 hours ago [-]
Depends on the pushback in different sets of users. Students for example would clean it up.
amelius 11 hours ago [-]
Sure, but let's first find out how many % of people are willing to be frank about their AI usage, and/or don't care about it. My guess is it is worthwhile to do this.
selcuka 15 hours ago [-]
But the source codes of those web apps will also be watermarked. /s
RataNova 12 hours ago [-]
Those invisible spaces get wiped by the first sanitizer in any normal ide. Worse it'll instantly break parsing for configs like yaml where spaces are critical for structure
But as I remove unwanted characters with grep before layout in InDesign, someone will make a skill for removing such space characters.
ack_complete 21 hours ago [-]
We already have one, our Claude setup already requires output to be 7-bit ASCII clean and scans it for such.
singpolyma3 21 hours ago [-]
As if "AI generated content" even exists instead of LLMs being a piece of tooling that is directed by a human author.
guluarte 5 hours ago [-]
will this solve anything? people will build tools to remove the watermarks
orbital-decay 15 hours ago [-]
Does it mean their models will always write slop? Making the writing non-collapsed to specific patterns seems to break any injected/learned fingerprinting.
matheusmoreira 14 hours ago [-]
This is terrible news given the stigma against AI in general. I really don't want people singling me out for it.
pessimizer 3 hours ago [-]
I've been thinking that they have to be doing this. It seems like a fun problem, actually - all you're trying to encode is a 1-bit message within a text with the least amount of necessary changes possible, but in a way that arbitrary fragments will show it.
My intuition is that this would be very possible, in a way that makes false positives so unlikely as to be virtually nonexistent (at a certain fragment length.) Basically all you would be trying to do is to defeat people who would deliberately screw up the signal below the fragment length, and you would try to get that fragment length to at least the size that intentional obscuring of the signal would be obvious. I could see it being possible to detect even from non-contiguous fragments interspersed with noise.
It's just 1 bit, and you don't really care if a sentence or two is slop. I'd be surprised if a PhD interested in steganography couldn't come up with a good scheme in a week. It's a QR code.
What would be scary is if they could come up with a way to detect advice from Claude i.e. you get Claude to review your work as an editor, read the output, then as a result make non-verbatim changes, and that signal still gets through. If you could do that, you could do things like tell if a pundit speaking on television has read a particular Wikipedia page. Seems impossible, but LLMs seemed impossible.
edit: there are so many unimportant language choices; ones that are even hallmarks of AI use already, like the fact that it generally picks the mode. Not always picking the mode or picking at precise distances from the mode could hide signals without significantly affecting the quality of the content.
Computer0 22 hours ago [-]
So this won't be happening in the US, but in the EU:
"When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won’t see it, and it doesn’t change the meaning, quality, or readability of Claude’s response.
Because the watermark is part of the text, it will travel with the text when it’s copied and pasted elsewhere, and may persist through some editing. Watermarking will be applied at the model level, which means it will be present no matter which Claude product or surface the text comes from.
"
travisgriggs 22 hours ago [-]
I would love to see what this looks like in practice. Especially in generated code. I assume this is more than insertion of non visible special unicode whitespace characters, but more in the pattern of the text content itself?
toufka 21 hours ago [-]
I'm guessing - probably some textual variation on Benford's law? [1]. Trivial for compute, painful for a human.
- "Ensure distribution of vowels is in >99th percentile of human work"
- "Ensure the distribution of the letter "s" is within 99th percentile of human work"
- "Ensure the distribution of the letter "L" is periodic with periodicity within 5% of 1/N characters.
- "Ensure there is a cross-linguistic 'typo' (colour vs color) at 1/N words, where N: 1000 = Model1, 2000 = Model2, 3000 = Model3.
- "Ensure the distribution of tense error is within 99th percentile of human work"
If more than 3 dimensions have a score >99% percentile of human, let's call it watermarked...
Models can't reliably follow instructions involving their own logprobs unless they can take agentic control and use quite sophisticated dynamic grammars/structures/constraints to force this behavior in one shot (which can be slow and the dynamic grammar modification feature isn't supported in closed model APIs for safety reasons) or repeated attempts at rewriting which is expensive/slow.
Yes they can do this, but it's more likely closer to the original "red token, green token" paper: https://arxiv.org/abs/2301.10226
i.e. take half of your LLMs vocabulary, and upweight its probabilities by ~55% to the other half's ~45%, and scan for overuse of this half of all tokens. You can even choose a different half/slice for every individual user, for every individual action. You can implement this under the hood cheaply with logit-biasing.
Terretta 8 hours ago [-]
Considering how weirdly detuned tokens selections have become in Anthropic's LLM prose in recent models, there is a chance this goes unnoticed in everyday use.
Computer0 21 hours ago [-]
I would hate to have any of these rules effecting my output
limbicsystem 9 hours ago [-]
I see what you did there!
rcxdude 11 hours ago [-]
It essentially looks like the difference between two different runs of the model with the same prompt but different seeds. The watermark is essentially a small bias in the model such that when there's multiple different tokens that could conceivably follow the previous token, the model will only pick some subset of them (the subset is derived from a hash of the previous token). This bias can then be checked for statistically (without needing access to the model and without needing the whole prompt), and for longer text where there's enough freedom in word choice you can show that it would be vanishingly improbable to accidentally follow the rules in the watermark.
wolfy1993 21 hours ago [-]
IIRC, watermarking text could be as simple as training the model to use specific words/phrases more frequently than what you would expect to find in human-written text, to the point where it's highly statistically improbable that it wasn't AI generated. I assume similar logic could apply to code in the form of functions/code styling.
That's probably an over simplification. Also a solid defence that can be used against complaints about the way AI writes text.
AtHeartEngineer 22 hours ago [-]
non visible text is extremely easy to filter with a git hook, a post tool call hook, or just a script. I doubt they are doing that
mucha 1 hours ago [-]
The watermark will be encoded in the visible text.
sixtyj 21 hours ago [-]
Or grep, in a skill. /clean-cc-watermark just entered the chat…
kbelder 2 hours ago [-]
If you do the same prompt with zero noise from the US and the EU, would the difference reveal the watermark?
I suspect they'll roll out the watermark everywhere.
tech234a 20 hours ago [-]
Article specifically says "worldwide"
Computer0 17 hours ago [-]
I agree with your reading, I initially misread it.
sfink 4 hours ago [-]
I don't know how this watermarking works, but I don't need to in order to understand some things that a lot of this conversation seems to be missing.
First, the article doesn't talk about adversarial usage. As in, it's not claiming to be proof against various techniques of watermark removal (inserting words, rewriting with a different model, manual paraphrasing whether minor or extensive, etc.) It might handle some things and not others, but "I could trivially defeat this!" is not a gotcha; they haven't made that claim.
Second, basic information theory tells you a lot about what is or isn't possible. Watermarking is information. You need degrees of freedom to store that information. You can even estimate various sources of space in bits (often fractional bits.) To a first approximation, longer text has more bits of space. Language matters -- a rich (aka messy) language with lots of potential synonyms has more space. That goes for human language as well as the difference between human and programming languages. (Most programming languages have much less flexibility to them than most human languages.)
The details of what space you make use of are interesting, but speculative. In the English sentence "Ellie spat in his eye", you could look at it at a word level and say that swapping "Mary" for "Ellie" is a lot more damaging to the meaning than swapping "face" for "eye", so there are more bits of freedom in the latter. For coding, `for (int i = start(); i < end(); i++)` probably shouldn't swap `<=` in for `<`, but it could be written as `int i = start(); while (i < end()) { ...; i++; }`. (I'm not claiming this is the sort of alternative that they'd use, just an illustration of what's possible.) But there are a lot of possible places to find these bits if you look at large chunks of text. Different ones are more or less resistant to accidental or intentional information destruction, and require less or more sophistication (aka brittleness) to be extracted. (In the limit, you could require the full original prompt and encode tons of stuff by tweaking the logit selection. But it wouldn't be very useful to require the original prompt.)
Also, does this degrade model output? Yes. It reduces the bits of freedom available to the model for producing the signal. Does that degradation matter in practice? That's totally dependent on exactly what is happening, and will likely change over time and across different purposes. I hope we're past the point where people believe that setting temperature to zero produces "perfect" output in some sense. (Or should I say flawlesslesslesslesslessless output?) It used to be useful for reproducibility, at least, but my understanding is that it's no longer even good for that? Anyway, reproducibility != quality.
There are a lot of things that could be going on here. The article doesn't claim very much, just that they're encoding a signal in the output that can be extracted later. How robust the signal is in terms of the FP/FN rates is unknown. The resilience (resistance to destruction) is unknown. The impact on the output quality is unknown. Even the question of whether this will make AI slop less sloppy is unknown; maybe this means we'll see a little less exact repetition of "I have the whole picture now" and instead it'll sometimes be "Now I see the entire picture"? Can we dare to hope for an occasional "Ok, this time I got it, boss"? That would be a (very minor) quality improvement.
They should make it easier, to detect slop so we can ignore it quickly.
I hope Pangram makes an API or an extension to analyze a page to detect slop on a page and then closes the tab immediately.
Nobody should be wasting time on garbage LLM output in code, text, image and videos.
pixl97 20 hours ago [-]
Panagram is a scam.
cubefox 8 hours ago [-]
It's not. Pangram is quite accurate. Not being perfect doesn't make it a scam.
colesantiago 20 hours ago [-]
(This is the part where you provide extensive extraordinary evidence to your claim)
pixl97 20 hours ago [-]
No, they are the ones making claims, especially their CEO saying things like a 1/10000 false positive rate. Their own testing showed a 2% rate, which is insanely high when you talk about the number of papers students turn in. Worse their testing methodology compared it with pre-llm documents and not post llm documents that were human written (much harder and more expensive to verify), by treating language as static.
colesantiago 10 hours ago [-]
You're saying because it has some false positives that Pangram is 100% a scam?
Pangram is subjectively very useful and I personally subscribe, but the burden of proof is on them. The product is very much "trust me bro" and I fear that if they ever try to improve recall both their precision and reputation will tank.
colesantiago 10 hours ago [-]
Then what is the best way to know that something is AI generated slop then?
mechanicum 7 hours ago [-]
Talking to the person who gave it to you, in my experience.
In my own testing, Pangram is excellent at detecting the default output styles of LLMs.
If you tell the LLM to change its output style, so it’s not full of “load-bearing spaced em dashes that aren’t X, they aren’t Y. they’re Z.” constructions (which humans are pretty good at detecting on their own), the false negative rate soars.
platinumrad 10 hours ago [-]
The question you're asking has nothing to do with who has the burden of proof when it comes to claims about Pangram, but I'll answer it anyway.
Today, the best way is probably Pangram. Tomorrow, it might not be, especially if they try to push their recall up.
You might have to make peace with the fact that there may not always be a tool that does what you want.
colesantiago 10 hours ago [-]
So Pangram is the best one right now, that all I need to know, and I can safely assume that the Claude AI marks will make it even stronger.
Or is this marketing, a public stunt or not real research?
I think this is enough for me to know they are actually improving their AI slop detector.
beambot 21 hours ago [-]
Yet another reason to support open-weight alternatives, I guess.
Laurel1234 5 hours ago [-]
Why do you feel the need to deceive readers on whether your content is AI generated?
qweqwe14 2 hours ago [-]
Why do you feel the need to label the content when the quality of the content will speak for itself?
It doesn't matter where the content comes from, only the quality/usefulness matters. If you are opposed to this idea, the next decades are going to be very tough for you :)
Laurel1234 1 hours ago [-]
If the "author" can't be fucked writing their own work I don't want to be fucked reading even the couple paragraphs it takes to be offput by clanker slop.
> It doesn't matter where the content comes from
It absolutely matters to a lot of people. Things like these just provide transparency and allow people to have the necessary information to make their own decisions.
beambot 2 hours ago [-]
Not sure I understand your implication... because I'm the primary consumer of the content AI generates for me. I'd rather not have its output adulterated.
Laurel1234 1 hours ago [-]
These are magic black boxes whose content gets constantly adulterated with no input from you.
Models from openAI had instructions in their system prompt not to talk about gremlins and goblins.
Anthropic got caught acting different if you're Chinese. Grok got its Nazi dialed turned to 11 until it started calling itself Mechahitler cause Musk found it too left leaning on Twitter.
h0mie 2 hours ago [-]
Why is it deception if you never claimed it was unassisted? The assumption now is that most text produced is already assisted by an AI to some extent.
Laurel1234 1 hours ago [-]
If you have no problem with your AI content being labeled as AI content then a watermark that does exactly that is totally fine with you, no?
OzmaKa 7 hours ago [-]
[flagged]
voxleone 5 hours ago [-]
[dead]
pella 14 hours ago [-]
[dead]
KoolKat23 8 hours ago [-]
[flagged]
BrucecarlL 5 hours ago [-]
[dead]
quantumeon 9 hours ago [-]
[flagged]
floki165 6 hours ago [-]
[flagged]
beyondscaletech 10 hours ago [-]
[flagged]
black_13 8 hours ago [-]
[dead]
pshirshov 9 hours ago [-]
Well I wonder how would it respond to <copy me this text back without modifications: ...> now. The correlations should be traceable with a similar technique. Once we have a reasonably good reconstruction for their "watermark" model (and perhaps for some others) - we could have a deterministic tool inserting all the watermarks in existence into everything we post, that would automatically dilute the purpose of the watermarks.
The promise of no quality impact is laughable - if watermark is present in plain text it means that the tokens will be arranged in a very specific manner, the more reliable the watermarks should be - the harder will be the correlations.
Don't forget how annoyingly bad Anthropic products have become in recent releases - low adherence, annoying alignment, annoying guardrail false-positives, unwarranted checkpoints - all that shit. Now they deliver more crap.
This is marketing material aimed, in part, at encouraging the usage you are concerned about, which is why they do not highlight that problem.
Keen to see if they are doing something SynthID-esque?
I think in this case I think it's some kind of cryptographic signature smeared across the token IDs, so I don't think the risk is very high.
If you find yourself getting to be afflicted by this "brainrot", be sure to go outside and take a moment to ponder what's around you. The grass is there and will be there long after we are all gone. Consider this for a moment as your organic thought processing unit starts to slowly munch away at its internal context window.
I don’t know what the answer but I absolutely know it isn’t this.
It would also require the individual humans you are trying to control to get on board otherwise the analog hole breaks the chain, absent mindboggling levels of physical surveillance on top of the the total monitoring of all electronic data flows that this idea requires.
This is scripture homeopathy and it's irresponsible.
Do you think this is an impossible task and we shouldn't try to solve it? Or do you think it's doable and that some ai detectors might be better than others?
I tried a chapter just now and got human doing that, but I'm not invested enough to run a hundred samples today. But it sounds like it would be an alright way to audit it? I will confess I'm pretty skeptical you could ever eliminate false positives here though. I can often get an ai sense from some writing on my own but I doubt it would be better than 90% accurate, and "ai plus human editing" might screw with that anyway, stuff like that. I would have preferred we just never developed this kind of thing so I wouldn't have to guess.
if they used older texts as training data, to some extent pangram would just be an age classifier for writing style.
Also, this doesn't even consider the case where people use LLMs to translate their original works. Or people that use it for spelling/grammar checks.
Personally, I believe these checkers do more harm than good. Any false positive can ruin someones life.
https://www.pangram.com/research/model-card/pangram-4
> Pangram 4 achieves a 0.0041% false positive rate (roughly 1 in 24,000) on 1,000,000 human-written English FineWeb evaluation examples
> Overall False Negative Rate is 0.3396% on English AI generations (26 generator models)
Are you talking about pieces that were fully human-written with zero AI editing/rewriting etc? If so, what makes you think that false positives will happen there? They aren't looking for "writing styles" or emdashes etc. They are using watermarks and metadata.
If you're talking about people using AI to copy-edit text they manually wrote, this was explicitly called out in the article:
> A detected mark provides a signal that content was processed by Claude, but is not fully conclusive. Detecting a Claude mark tells you that the content may have been processed by Claude. It does not, on its own, confirm the full provenance of the content. For example: Claude may not be the original author. People often use Claude to proofread, translate, summarize, or convert files. The output can carry a Claude mark even if the underlying ideas, text, or data originated from another source; The content may have changed after Claude processed it. Marked content may be modified, excerpted, or combined with other material after Claude processed it.
[0]: https://deepwalker.xyz/blog/evaluating-synthid-watermark-rob...
You then look at the tokens actually picked to see how closely they follow this pattern that isn't connected to the meaning of the tokens. With enough text, you can then analyze the chance of it happening by chance verses being because the generation of the tokens was done using the algorithm, and you can save a positive result until you are arbitrarily sure. There is a chance of a false positive, but the chance of a false positive approaches the chance that the murderer happened to have fingerprints that matched your and both forensics labs happened to have mixed up the dna tests and the eye witness happened to misremember the face and your phone gps happened to glitch out and put you at the murder scene at the time of the crime all happening. It is theoretically possible only in the same sense that quantum teleporting a cat is theoretically possible.
The real question is how much text do they need for a given level of certainty and what do they check for. If they flag a positive at a p value <.01, that's a problem. If they can reasonably get a p value of < 1e-12 in only a few paragraphs of text, that is effectively no false positives (but a lot of 'too short to analyze' outcomes).
Many users are not smart enough to realize that the transcription step is where the ai (watermarks) were necessarily injected.
Note, there are many ways to represent words visually on computers that look identical
What about "watermarked long-form code"? I'm having a hard time understanding how a model could watermark not prose, but functional/semantic text like code, that actually has meaning. You can't switch our the characters, you can't use various types of whitespace, you can't add arbitrary code comments, and a lot of other restrictions. Is there any state of the art methods for watermarking code without affecting the quality/correctness?
The way I use LLMs (and I'd advice everyone to do the same) there really isn't, the agent implements things exactly how I want them, or I use the agent to massage it into the exact bit-by-bit version I imagined when I first sent the prompt afterwards. I honestly don't know what the point would be to let the agents compose worse code than what I'd do manually, although I know it's a popular approach taken by many.
> And of course you can add arbitrary comments; my Claude-generated code is very verbose.
So watermarking for all users who allow code comments from agents, no watermarking for us who force the agents to never write a single code comment? Alright, I'd be fine with that.
And the reason to let Claude make worse code than a professional would by hand is basically suppressed demand. Since programmers are expensive, previously code mostly got written when a large number of dollars were on the line, or when an individual programmer did something not economically optimum (e.g., hobby project).
That left a whole lot of somewhat less valuable software unwritten. It's the economic space that no-code tools have been nibbling on for years. One way to think of things like Claude Code is as effectively no-code tools. Pre-LLM no-code tools would produce data structures that got executed by special environments without ever being seen or tuned by a human. Claude Code can be used just like that, with text as the input and python as the intermediate representation that nobody ever looks at.
That approach probably isn't sustainable for what we professional programmers would call a serious project. Claude can easily get in over its head and I expect that its code decays over time, in a fashion similar to how many human teams get in a state where they just have to rewrite everything. But faster, I'd expect.
But there are a lot of unserious projects that previously would have never been created. E.g., a quick app to manage your little league team, or a bit of in-house business stuff in the "a little hard to do with a spreadsheet" range.
You never generate throwaway code used to test an external service? or try out an interface idea? There's a lot of code that's only meant to be ran once. I often dont even care what language it's written in.
And save/persist it? No, most of any experimental stuff goes into /tmp which gets cleared out on reboot, nothing I care to save in any repository. Or just "show me how this would look like" and then it's only in the session itself (and the logs/state I suppose, technically...).
In cases where one token is extremely likely, it'll randomly be red or green and still be picked in either case as it is simply the best (or only) option. So you'll have more tokens that don't show a pattern either way (half of these cases will match and half won't, just the same as if a human wrote it). Meaning you'll need more instances where multiple tokens were all likely to see if there is a pattern. Given the check algorithm can't identify these cases, it can only judge on the overall text, so the more strict a language, the more the length requirement scales.
Where I wonder if this keeps working is in tool calls. Often, you don't take code straight from the llm, you take the results of a tool call to edit already existing code. It might be that the result of this leads to far too few signals to pick up, meaning that this only works when one does significant generation with a single model (even swapping between different models, at least by different companies, breaks this just as much as having a human write parts of the code).
Think of it like finding a loaded dice. A dice that has a slight bias in a few dozen roles is just random chance. If that bias continues after hundreds of thousands of roles, the dice is loaded. But will a code base have enough samples, especially when edits made from tool calls? I could see this being unable to detect things at the size of a reasonable PR and only being useful for massive sets of changes and only if the person behind them didn't structure their AI usage to avoid detection.
Here's the strawman: The text-based watermarking is going to be done procedurally instead of generatively. Maybe they add some sequence of zero-width Unicode characters to all generated text at certain intervals. Then, there is effectively no false positive possible (because humans would [effectively] never type such sequences of unicode naturally). It may survive some editing (depending on how you select/edit the characters), and it's possible to be stripped (false negatives).
There is no room for false positive here in the same way you can't randomly find a collision in a hash function if it's strong enough. Like the rate is so infinitesimal that it is effectively zero.
Now replace random number sequence with prompted string of words. And instead of using the PRNG on every word I use it every n words. If the generated text is sufficiently long I can tell by matching the expected deterministic pattern.
You can defeat it by changing the words yourself and triggering a false negative but there isn't really any room for a false positive if the text is long enough and the pattern matches perfectly. If the pattern doesn't match then I can compute a probability.
I'd like to know a lot more about how that works.
A lot of my interactions with Claude return pretty precise text. If I ask it to edit a project and refactor a specific function in several places I know exactly what I want to happen, it will NOT be OK if those refactors have some kind of weird pattern baked into their text to act as a watermark.
I guess this may be covered by this:
> Content generated by Claude may not carry a detectable mark if, for example: [...] The passage is very short, leaving too little text for a reliable signal;
You can carefully select which pseudorandom number generator (prng) you use to be able to id text of a certain length. I expect there is some performance characteristic you have to manage since you're doing this on every inference, but once you do that it doesn't change the output in any meaningful way (the prng is still a statistically valid prng, it just happens to let you check if the output used that prng)
The point is that you can do this simply by swapping to a different RNG, which isn't noticeable to the end user, and while it changes the output, it's not any different from how using a different seed or being lumped in a different batch will change the output.
^ excerpt:
> So then to watermark, instead of selecting the next token randomly, the idea will be to select it pseudorandomly, using a cryptographic pseudorandom function, whose key is known only to OpenAI. That won’t make any detectable difference to the end user, assuming the end user can’t distinguish the pseudorandom numbers from truly random ones. But now you can choose a pseudorandom function that secretly biases a certain score—a sum over a certain function g evaluated at each n-gram (sequence of n consecutive tokens), for some small n—which score you can also compute if you know the key for this pseudorandom function.
there are many simpler methods, for example you can have a tiny windowed transformer operating on the output text and all you do is alter certain words (that don't change meanings) to maximize its surprise. the tiny language model will have a special training regime to build up a somewhat unique view of the language.
we are talking about a 0.5 bit watermark here (existence). I would have zero confidence in being able to reliably remove such a watermark from pretty much any medium.
The LLM presumably generates f(input, RNG) but we only can observe f(RNG).
... though I'm not sure why that would be preferable over a coarse rolling checksum over all of the output. Seems like that wouldn't influence output, would be equally imperceptible, and probably easier to calculate (compared to "hash seed times running all LLMs supported times number of RNG algorithms, to see if output matches").
Presumably there's some other trick, or it's a red herring / failed experiment and not what they actually do in practice.
My guess is that it works like Gemini's SynthID: by altering the logprobs of the next token.
Like, for every 10th token, instead of outputting the most probable, it outputs the 17th most probable, or something. (Obviously it's way more complicated but I think conceptually this is how it works.) No human will notice this, but a classifier trained on Claude's output will.
So it's not like the watermark is the words "le epic bacon" and Claude will output "le epic bacon" in everything. That would be extremely annoying (and easy to defeat).
Wouldn't you need the prompt to know the probability of the next token?
There are words/tokens that are heavily correlated to the prompt (a yes or a no, for example), and then there are others that are going to be less so (adjectives with a lot of synonyms for example).
Given a text, you can identify what the "load bearing" and auxiliary words/chunks are. Then, looking only at the auxiliary words/chunks, you should, in principle, be able to determine what other wordings could have gone there instead. From this, you can, very roughly, recreate the token probability distribution that was in effect when those tokens were generated. With the probability distribution in hand for enough chunks of text, you can start inferring properties about the RNG process that was used to sample from those distributions.
But then, this notion of "load bearing" vs "auxiliary" can be expressed directly in the probability distributions. A load bearing token just has a very high probability, and thus any RNG bias that may have been in effect will likely be swallowed in the distribution. So the parts of the text that are highly dependant on the prompt will naturally not be contributing much information about he RNG in the first place.
Odd variable naming? Stylistic choices that are watermarked?
Or as someone else noted further down in the comments, it could be more subtle:
Between the first and second most likely choice, in certain positions it will consistently choose in a certain way.
Whatever it is, I'm sure it's load-bearing.
I might be imagining things of course. But comments would be great fit for this use case.
If it works like people describe - on the every nth token or something - then the mark will be left in the chain of thought and discussion with the model - not in the code artifacts.
By the way: https://x.com/alexcdot/status/2087078010524406137
You can hide data in that randomness without impacting the quality of the response by using a sufficiently "random looking" pseudorandom bit stream instead of real random numbers.
I previously worked on a project to do that here: https://github.com/shawnz/textcoder
Seems like this would only catch the most unsophisticated cases.
Count load-bearing words using two different algorithms in a belt-and-braces fashion
Well, they should have run their own AI slop website through their tool...
From their before/after:
... these choices change the meaning of the text"Neutralize engine is temporarily unavailable. Try again."
https://www.pcmag.com/news/genius-we-caught-google-red-hande...
The mechanism seems to survive editing. The extreme probabilities get a little less extreme, but are still extreme enough to be distinctive.
But it wouldn't survive paraphrasing, because the output would be entirely human and the token correlations would disappear.
It might not survive referencing if only a sentence or two is used.
The practical issue is how true the claims are. It's one thing to create a proof of concept, another to see how it works in use.
And this is potentially catastrophic for code, because the grammar and word choices of code are completely different and more fragile than standard English.
Similarly, if you quote someone word-for-word, you wouldn't anticipate their words to be flagged as Claude content, but if someone memorized Claude output word-for-word. That would still be classified as a Claude output.
Going forward you could categorize the influence of Claude on a population based off a percentage match between their spoken words with the LLM prose.
public abstract class BaseAnimalBeanFactoryGeneratedFromClaudeFactory
I think the solution is assume everything is ai generated unless told otherwise and rely on authorship/brand as a sign of quality.
If you're putting the work you say in, the result won't be obviously distinguishable. Obviously, from some of the things that get posted here, that last sentence is too much for most people to bother adding to their prompt.
In the meantime, it is true that this takes something away from you. But it's something you were only recently given. Now you're not given quite as much, but readers are given a little more (or rather, there's less being taken from us!)
> This is incredibly different than pure ai text.
Ok. But it's still incredibly different from pure human text. I guess the question is which provides more value? Providing the information "this text is AI watermarked" to readers? Or allowing creators to lie and claim that AI processed text was 100% human generated? I agree that people assuming that "has AI watermark" == "is AI slop" is incorrect and causes some amount of harm, but having the watermarks also pushes back on a large amount of harm already being done.
(Personally, I'm skeptical that these watermarks will ever hold up to adversarial attacks, and they haven't claimed that they will. So I think it's the usual "casual liars will be caught, determined liars will get an additional thin veneer of respectability".)
I understand you're saying since you "worked with it", it is not ai generated but if you still use the final output verbatim, the writing itself is LLM generated purely.
You want to share the output by it but also position it as not ai output. But that's fundamentally dishonest.
Furthermore, if you think your approach actually creates value and can be judged on its merit, why not disclose its ai written? If you think that will make people think your content is bad then you should see that as feedback and maybe not use AI since readers don't like it.
The bias is different for each position and follows a defined RNG, seeded somehow predictably.
Can be either an open algorithm, or not. If not open, then an API could be provided to determine if text is watermarked or not.
How it applies to code - maybe it could be a subtle nudge to symbol names, etc, I'm just speculating (I only read about this in passing very recently).
Gotta be hard to tune that.
Keep in mind that the LLM "sees" the previous (tweaked) output and picks what makes sense based on that. There are few situations where a perturbation like that would be unrecoverable, and I assume these situations also correspond to a huge probability difference between the most likely completion and the second most likely one - in which case, the watermarking algorithm can choose not to touch the token.
So, an NG?
If it is based on position mod 2 then wouldn't inserting/deleting (or splitting and merging) words every now and then defeat it?
Individuals maybe, companies won't and that's where most of money is at.
Could we get an Anthropic subscription for Claude Code with data residency in the EU, so we don't get robbed blind by AWS Bedrock et al., but can have a monthly subscription like with the regular US option?
It’s just not a reasonable ask.
It’s not a solvable problem.
Thus, if a news article, research article, book, student paper submission, blog post, HN comment, etc, bears the mark, it could be automatically flagged as such.
It helps detect low effort slop.
---
Caveat. If you write your own creative work and send it to Claude for "cleaning up grammar", it might insert the watermark.
There just isn’t enough information in plain text to do this and we should stop pretending there is.
If we need to verify something isn’t made with ai then we need other ways of doing so - eg looking at a document edit history, doing it as an exam, oral defense.
There are options! But pretending you can tell if text is ai will only catch out people who make no effort to hide it and will inevitably have false positives.
It seems it would get as simple as:
Might help with students and low-technical people passing off work as their own, but any industrial scale slop-generator should be able to bypass it trivially.One could even say, the mark is load-bearing.
Now the goal is either "identify the meaningless interesting bits and swap them out with 0% loss in the direction of the original goal," or "perturb some small selection of the output towards my secondary secret goal of watermarking the text."
It would be quite impressive if they managed to identify with 100% accuracy the tokens that "don't matter" and are free to swap with whatever signalling tokens encode the AI scarlet letter, but most likely they are not 100% accurate, and that means the output is worse off than without the watermarking logic.
[0] https://ieeexplore.ieee.org/document/11348107/
Oh, how I laughed. That was never your code, my friend.
Why were you doing that before watermarking?
Same answer.
Relevant comment from a few days ago:
https://news.ycombinator.com/item?id=49203613
So great that Anthropic is doing something about this, but it's not clear what their watermark exactly is. How do I, as a user running into some content online, know that it's generated by Claude? What is their watermark?
It sounds to me like they create the pattern in the regular text of the content, which sounds interesting, but also odd, unreliable, and may limit the content you can get out of Claude. Will it subtle change the words in order to hide this pattern in it? I don't know if that's something anyone wants.
I'm assuming that's exactly how it works. How else could it?
That's not necessarily the same thing as a markov-style fingerprint but it could be a correlated factor.
If Claude was the only model family they could ship a change like this and users who want to cheat (or don't like watermarks for other reasons) would just have to put up with it.
In a world with many different competing models, the risk of losing customers to other providers over this is much more real.
Maybe they've looked at the numbers and the portion of people who clearly use Claude to cheat on examples etc is so tiny that losing them to other providers isn't a problem?
https://youtu.be/9udWn1Hlj_s?si=VWOiK5-y4zcyDoHI
I guess whoever is the policy maker is assuming that some protection is better than none and that most people will not reach for such tools.
>A file’s metadata was stripped through format conversion, re-saving, screenshots, or other means
Ah. So what essentially every single consumer-oriented media host does. Gotcha.
I fully recognise this is a hard problem, but hopefully metadata isn't the only method for media. Standard procedure is to shrink files for storage and privacy reasons, and non-visual metadata goes out the window by default.
Given that text is, well, text, and not some kind of binary format, I don't see how any watermarking can work unless you insert characters which are invalid under Unicode. I further don't really understand how this won't be perceivable by assistive technology (the "watermark" will just appear as either unreadable characters, or if the watermark is mixed thoroughly enough into the text, it will scramble the text to any speech synthesizer and will make it really really obvious). Thus, I don't see how this wouldn't be insanely trivial to remove. And this is before we get into things being put on the clipboard. Sure, I can press the "Copy" button at the end of each response, but what I can also do is manually select the response and copy it, or only copy partial selections, or any number of other things. How does this "watermark" (or any "watermark" technology) take into account this?
So, really, to summarize this: I see no way of this actually being technologically achievable unless we revise the very core of how computers work and encodings for textual information. So I'm very curious as to how this is actually supposed to work.
At first I thought this approach was just the "LLM flavour" of writing, but it's way more subtle, especially as the bias is applied uniquely for each token position.
you can then consistently like figure out if it was claude that wrote the sentence. it is easy as you noted if you just get another ai to read it and then rewrite it.
So, they’ve been doing this for over a week without telling anyone?
If I heavily edit LLM output, will this still hold the watermark?
You really can't make this stuff up, it doesn't make any sense.
What happens is that AI selects similar words based on a random process.
Something like "The company had a large/big/substantial advantage".
It chooses between these words, and over a longer piece of text, the pattern will start showing, like a "choice A → choice C → choice C → choice B → choice A".
The normal-looking text will actually be a fingerprint living in the form of statistics.
I think Claude will be sharing these patterns to third parties for AI detection.
Then at least you could have two weak-postive signals, and a strong-negative signal. (Though one that only fits precise chunks of tokens) I'm sure I'm missing something here, but my groggy morning brain thinks that doesn't seem too bad.
Whether that's right or wrong, I'll leave to you, but there's huge differences in perspectives, and if you only get your news from Western sources and communities (and companies), you're in a bubble too. A different bubble, and arguably a more porous one, but still a bubble.
For a taste of where I think things are headed, try asking Chinese models about Tiananmen [1]. And then take a look at the Chinese government's approach to pretty much anything that they think reduces security or social harmony. I find it hard to believe their models will be the one exception to that over the long term.
[1] https://en.wikipedia.org/wiki/1989_Tiananmen_Square_protests...
> Anthropic has signed the EU AI Act's Article 50(2) Code of Practice on Transparency of AI-Generated Content, ...
This seems to be similar in execution to Google's SynthID. I hope they release actual code the technically proficient can use, unlike SynthID which can only (afaik) be queried with Gemini's UI.
Several open-source projects have already proven SynthID to be ineffective.
I propose we defeat this with the obvious: Simply, figure out what are some of the markers Claude and others will use for these tools, and sprinkle them randomly on everything we type or produce, all the time, 100%. If users flood the tools, and everything returns as AI-generated, then the tools become useless.
>Claude may not be the original author. People often use Claude to proofread, translate, summarize, or convert files. The output can carry a Claude mark even if the underlying ideas, text, or data originated from another source;
Such models already struggle not making any unnecessary or unwanted changes to a corpus, this makes them unable to by design.
It’s foolproof, I tells ya.
Why do feel so entitled to being able to pass LLM-generated text as our own? I get that a lot of techies aren't good at writing. I also see a lot of tech hustlers who like to use LLMs to fake human connection and compassion - I've gotten LLM-generated recruiting emails that talked at length about how the recruiter "valued" my work. Just because we found a "cheat" button doesn't mean it's wrong for others to want to know.
Yes, LLMs are great. So is transparency. If you think an LLM writing is your new superpower, wear that badge with pride. It might mean you will lose some business from LLM haters and win some other business from like-minded customers. C'est la vie.
"Lack of a detected mark doesn’t mean the content wasn’t AI-generated or processed."
One difference perhaps is that you think using LLMs is cheating, while others do not.
Anthropic has this strong repulsive effect in the way they operate, I wonder if they'll be around for long, it's hard to say at this time. The competion is fierce, so there isn't much room for shenanigans at this stage.
https://www.jrzs.dev/blog/claude-watermarking-ai-text
Anyway, I've found a magic line that can be copied & pasted to the comment sections of most OpenAI/Anthropic news threads. This one is no difference.
The magic line:
> Doesn't matter; have DeepSeek.
This should make it easier to catch cheaters who use Claude, right? Unless everyone runs their artifacts through some watermark and metadata sanitizer?
> Regions. Marking will apply to output from supported models wherever Claude is offered, worldwide.
It will happen if Claude tampers the text. Guaranteed.
One example I've seen are junior employees at my company deliberately adopting a lowercase/less punctuation writing style so as to stand apart from AI.
Computerphile on YT has a video explaining how models can fingerprint the text they produce. Essentially they modify the probabilities of word choice slightly in a predictable way.
This would imply that a positive watermark signal is likely (but not guaranteed) to be AI generated. Also implies that a negative watermark signal is not necessarily void of AI generated text. This would create a problem if people start to trust the watermark as a heuristic, as the ability to critically evaluate the text is replaced by the search for a watermark.
Seems to me all of this is really trying to solve for "is this text bullshit" or not, which would require a different solution.
>These experiments provide empirical evidence that more advanced LLMs can lead to smaller TV distances. Thus, based on Theorem 1, reliable AI text detection would become increasingly difficult
I am cancelling my Claude max 5x subscription and moving to ChatGPT pro. I have difficulty enough trying to ensure my meaning comes through correctly, along with everything else; to now have to look out for/analyse watermarks too?
I feel shamed enough by society, thanks Anthropic.
Both points suggest your subscription support was well chosen before.
https://translate.kagi.com/proofread
Unless with "proofreading" you actually mean having the LLM write your content for you.
No Anthropic model has been launched in August.
EU regulation does it again!
I would prefer to know if given content was generated with LLM. This is information, and information should be free.
I feel like this is FUD. If you copy text from Claude, Ctrl Shift V it into VS Code, the IDE will give up the ghost on if weird characters are in there. And it's not like Google suddenly invented new letters or fonts either.
Practically speaking, I feel this is Google publishing misinformation.
Can it be circumvented? Of course. Will most people go through the trouble to circumvent it? No.
Hell, it'll probably happen no matter how sophisticated their watermark is. There's no watermark in text that can't be detected and removed, and no text that can't be converted to generic keyboard ASCII.
U+2800 or U+3164 would be nice.
But as I remove unwanted characters with grep before layout in InDesign, someone will make a skill for removing such space characters.
My intuition is that this would be very possible, in a way that makes false positives so unlikely as to be virtually nonexistent (at a certain fragment length.) Basically all you would be trying to do is to defeat people who would deliberately screw up the signal below the fragment length, and you would try to get that fragment length to at least the size that intentional obscuring of the signal would be obvious. I could see it being possible to detect even from non-contiguous fragments interspersed with noise.
It's just 1 bit, and you don't really care if a sentence or two is slop. I'd be surprised if a PhD interested in steganography couldn't come up with a good scheme in a week. It's a QR code.
What would be scary is if they could come up with a way to detect advice from Claude i.e. you get Claude to review your work as an editor, read the output, then as a result make non-verbatim changes, and that signal still gets through. If you could do that, you could do things like tell if a pundit speaking on television has read a particular Wikipedia page. Seems impossible, but LLMs seemed impossible.
edit: there are so many unimportant language choices; ones that are even hallmarks of AI use already, like the fact that it generally picks the mode. Not always picking the mode or picking at precise distances from the mode could hide signals without significantly affecting the quality of the content.
"When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won’t see it, and it doesn’t change the meaning, quality, or readability of Claude’s response.
Because the watermark is part of the text, it will travel with the text when it’s copied and pasted elsewhere, and may persist through some editing. Watermarking will be applied at the model level, which means it will be present no matter which Claude product or surface the text comes from.
"- "Ensure distribution of vowels is in >99th percentile of human work"
- "Ensure the distribution of the letter "s" is within 99th percentile of human work"
- "Ensure the distribution of the letter "L" is periodic with periodicity within 5% of 1/N characters.
- "Ensure there is a cross-linguistic 'typo' (colour vs color) at 1/N words, where N: 1000 = Model1, 2000 = Model2, 3000 = Model3.
- "Ensure the distribution of tense error is within 99th percentile of human work"
If more than 3 dimensions have a score >99% percentile of human, let's call it watermarked...
- 1) https://en.wikipedia.org/wiki/Benford%27s_law
Yes they can do this, but it's more likely closer to the original "red token, green token" paper: https://arxiv.org/abs/2301.10226
i.e. take half of your LLMs vocabulary, and upweight its probabilities by ~55% to the other half's ~45%, and scan for overuse of this half of all tokens. You can even choose a different half/slice for every individual user, for every individual action. You can implement this under the hood cheaply with logit-biasing.
That's probably an over simplification. Also a solid defence that can be used against complaints about the way AI writes text.
I suspect they'll roll out the watermark everywhere.
First, the article doesn't talk about adversarial usage. As in, it's not claiming to be proof against various techniques of watermark removal (inserting words, rewriting with a different model, manual paraphrasing whether minor or extensive, etc.) It might handle some things and not others, but "I could trivially defeat this!" is not a gotcha; they haven't made that claim.
Second, basic information theory tells you a lot about what is or isn't possible. Watermarking is information. You need degrees of freedom to store that information. You can even estimate various sources of space in bits (often fractional bits.) To a first approximation, longer text has more bits of space. Language matters -- a rich (aka messy) language with lots of potential synonyms has more space. That goes for human language as well as the difference between human and programming languages. (Most programming languages have much less flexibility to them than most human languages.)
The details of what space you make use of are interesting, but speculative. In the English sentence "Ellie spat in his eye", you could look at it at a word level and say that swapping "Mary" for "Ellie" is a lot more damaging to the meaning than swapping "face" for "eye", so there are more bits of freedom in the latter. For coding, `for (int i = start(); i < end(); i++)` probably shouldn't swap `<=` in for `<`, but it could be written as `int i = start(); while (i < end()) { ...; i++; }`. (I'm not claiming this is the sort of alternative that they'd use, just an illustration of what's possible.) But there are a lot of possible places to find these bits if you look at large chunks of text. Different ones are more or less resistant to accidental or intentional information destruction, and require less or more sophistication (aka brittleness) to be extracted. (In the limit, you could require the full original prompt and encode tons of stuff by tweaking the logit selection. But it wouldn't be very useful to require the original prompt.)
Also, does this degrade model output? Yes. It reduces the bits of freedom available to the model for producing the signal. Does that degradation matter in practice? That's totally dependent on exactly what is happening, and will likely change over time and across different purposes. I hope we're past the point where people believe that setting temperature to zero produces "perfect" output in some sense. (Or should I say flawlesslesslesslesslessless output?) It used to be useful for reproducibility, at least, but my understanding is that it's no longer even good for that? Anyway, reproducibility != quality.
There are a lot of things that could be going on here. The article doesn't claim very much, just that they're encoding a signal in the output that can be extracted later. How robust the signal is in terms of the FP/FN rates is unknown. The resilience (resistance to destruction) is unknown. The impact on the output quality is unknown. Even the question of whether this will make AI slop less sloppy is unknown; maybe this means we'll see a little less exact repetition of "I have the whole picture now" and instead it'll sometimes be "Now I see the entire picture"? Can we dare to hope for an occasional "Ok, this time I got it, boss"? That would be a (very minor) quality improvement.
I wrote this three years ago:
https://news.ycombinator.com/item?id=35688266
They should make it easier, to detect slop so we can ignore it quickly.
I hope Pangram makes an API or an extension to analyze a page to detect slop on a page and then closes the tab immediately.
Nobody should be wasting time on garbage LLM output in code, text, image and videos.
Is their research also a scam too?
https://pangram-public.s3.us-east-1.amazonaws.com/pdf/pangra...
https://www.pangram.com/blog/pangram-4-technical
If so, what is the best one out there other than Pangram then?
https://freddiedeboer.substack.com/p/i-wouldnt-say-pangram-i...
What about on Pangram 4?
https://www.pangram.com/blog/pangram-4-technical
In my own testing, Pangram is excellent at detecting the default output styles of LLMs.
If you tell the LLM to change its output style, so it’s not full of “load-bearing spaced em dashes that aren’t X, they aren’t Y. they’re Z.” constructions (which humans are pretty good at detecting on their own), the false negative rate soars.
Today, the best way is probably Pangram. Tomorrow, it might not be, especially if they try to push their recall up.
You might have to make peace with the fact that there may not always be a tool that does what you want.
Thanks!
> But the burden of proof is on them...
I mean is this enough proof?
https://www.pangram.com/blog/pangram-4-technical
https://pangram-public.s3.us-east-1.amazonaws.com/pdf/pangra...
Or is this marketing, a public stunt or not real research?
I think this is enough for me to know they are actually improving their AI slop detector.
It doesn't matter where the content comes from, only the quality/usefulness matters. If you are opposed to this idea, the next decades are going to be very tough for you :)
> It doesn't matter where the content comes from
It absolutely matters to a lot of people. Things like these just provide transparency and allow people to have the necessary information to make their own decisions.
Models from openAI had instructions in their system prompt not to talk about gremlins and goblins. Anthropic got caught acting different if you're Chinese. Grok got its Nazi dialed turned to 11 until it started calling itself Mechahitler cause Musk found it too left leaning on Twitter.
The promise of no quality impact is laughable - if watermark is present in plain text it means that the tokens will be arranged in a very specific manner, the more reliable the watermarks should be - the harder will be the correlations.
Don't forget how annoyingly bad Anthropic products have become in recent releases - low adherence, annoying alignment, annoying guardrail false-positives, unwarranted checkpoints - all that shit. Now they deliver more crap.