앤트로픽이 클로드 워터마크의 작동 방식을 공식 문서로 내놨습니다. 8월 14일입니다.
사흘 전 보도로 알려졌던 내용과 갈리는 대목이 있습니다. 특히 '교정'이 그렇습니다.
표식이 붙는 자리와 붙지 않는 자리가 이 문서의 알맹이입니다.
■ 글자를 넣는 것이 아니다
앤트로픽은 요약에서 못 박았습니다. 텍스트에 무언가가 추가되지 않고, 숨은 문자도 없다는 것입니다.
그러면 표식은 어디에 남을까요? 단어를 고르는 순간입니다. 모델은 한 번에 한 단어씩 만들고, 다음 단어를 정할 때 후보 목록에서 하나를 고릅니다.
공식 문서가 든 예가 “오늘 날씨는 춥고…”입니다. 다음에 '달콤한'이 올 가능성은 거의 없지만, '흐린'이나 '잿빛'은 둘 다 자연스럽습니다. 어느 쪽을 골라도 상관없는 이런 선택이 한 글에서 수없이 일어나고, 워터마킹은 그 선택들에 패턴을 남깁니다.
독자는 알아챌 수 없고 키를 가진 쪽만 감지할 수 있습니다. 지폐나 문서에 찍는 워터마크와는 성격이 다릅니다. 눈에 보이라고 만든 것이 아닙니다.
그러면 앤트로픽이 새로 발명한 것일까요? 방식 자체는 새로 만든 것이 아닙니다. 구글 딥마인드가 2024년 네이처 논문으로 낸 신스아이디 텍스트의 한 버전이고, 2022년 스콧 애런슨의 제안까지 거슬러 올라가는 계열입니다. 토큰이 더 들지 않아 값도 오르지 않는다고 밝혔습니다.
■ 어디에 붙고 어디에 안 붙나
워터마크는 클로드가 고른 단어에만 붙습니다. 이 한 줄에서 나머지가 갈립니다.
사람이 쓴 글을 클로드에 맡겨 문법과 구두점만 고치게 하면, 표식이 붙을 수 있는 곳은 그 몇 안 되는 수정 부분뿐입니다. 앤트로픽은 그 양이 감지되기에는 너무 적을 수 있다고 적었습니다.
반대로 번역은 붙습니다. 번역문은 모든 단어를 클로드가 고른 것이기 때문입니다.
코드도 갈립니다. “2 + 2 =” 다음에는 '4' 말고 똑같이 좋은 답이 없습니다. 선택의 여지가 없는 자리에는 워터마크가 붙지 않습니다. 다만 주석처럼 표현을 고를 수 있는 곳에는 붙을 수 있고, 실제 코드에 미치는 영향은 미미하다고 했습니다.
사실을 다루는 문장에서도 성깁니다. “뉴턴의 가장 유명한 저작은 프린키피아…” 다음에 올 단어는 정해져 있습니다. 틀리면 안 되는 자리에는 고를 여지가 없습니다.
그러면 지우면 그만일까요? 지우는 것도 정도의 문제입니다. 가벼운 편집으로는 완전히 지워지지 않고, 모든 단어를 바꾸는 전면 재작성이면 지워진다고 했습니다. 다만 그 경우엔 그 글을 여전히 AI가 만든 것이라 부를 수 있는지가 논쟁거리라고 덧붙였습니다.
■ 그래서 무엇을 증명하나
워터마크가 답해주는 것은 하나입니다. 어느 시점에 클로드가 관여했을 가능성.
앤트로픽은 이것이 “클로드가 썼다”와 “클로드가 대폭 편집했다”를 구분하지 못한다고 명시했습니다. 사람이 썼다는 것을 확인해주지도 않고, 다른 AI가 썼는지도 알 수 없습니다.
추적도 안 됩니다. 워터마크에도 키에도 사용자와 조직, 대화를 알아낼 수 있는 정보가 들어 있지 않다고 밝혔습니다.
왜 하느냐는 EU AI법 때문입니다. 8월 2일부터 EU 시장에 서비스하는 AI 제공자는 AI가 만든 콘텐츠를 표시해야 합니다. 앤트로픽은 2026년 7월 EU의 AI 생성 콘텐츠 투명성 실천 강령에 서명했고, 서명 주체는 약 190곳입니다.
확인 도구는 아직입니다. 워터마크 탐지 API를 곧 내놓겠다고 했고 구현 세부는 정리 중이라고 했습니다. 이미지나 파일은 방식이 달라, png·jpg·svg 같은 형식에는 C2PA 라는 개방 표준으로 파일 메타데이터에 암호 서명된 메모를 붙입니다.
■ 클로드만의 일이 아니다
이건 클로드만의 일일까요? 표식을 넣는 쪽은 늘고 있습니다. 구글은 2024년 5월부터 제미나이 앱과 웹이 생성하는 텍스트에 신스아이디를 적용해 왔고, 오픈AI도 텍스트까지 생성물 출처 표시를 넓히는 작업을 진행 중이라고 뉴시스는 전했습니다.
앤트로픽은 앞으로 출시되는 모델부터 적용하고 기존 클로드 모델에는 수개월에 걸쳐 넣을 예정입니다.
다만 방향이 한쪽만은 아닙니다. 구글은 14일 이미지와 영상, 음악 생성물의 눈에 보이는 워터마크를 사용자가 켜고 끌 수 있게 바꿨습니다. 콘텐츠 안에 신호를 심는 신스아이디와 출처 정보를 담는 C2PA 는 그대로 두면서입니다. xAI 는 그록이 작성한 텍스트에 기계 판독이 가능한 비가시 워터마크를 적용한다는 계획을 아직 공개하지 않았습니다. 이미지와 영상에는 가시 표시를 적용하면서도, 텍스트는 이용자에게 ‘그록으로 작성’ 같은 문구를 표시하도록 권고하는 데 그치고 있다고 지디넷코리아는 전했습니다.
보이는 표시는 줄이고 기계가 읽는 신호는 남긴다. 업계가 공통으로 향하는 지점입니다.
■ 이 블로그가 고친 것
덧붙일 것이 있습니다. 이 블로그가 8월 11일에 쓴 글은 국내 보도를 근거로 “사람이 쓴 글을 교정하거나 번역만 해도 표식이 남을 수 있다”고 적었습니다.
공식 설명을 보면 번역은 맞지만 교정은 다릅니다. 문법만 손보는 정도라면 붙을 것이 거의 없다는 것이 앤트로픽의 설명입니다. 해당 기사는 이에 맞춰 고쳤습니다.
표식이 늘어나는 방향은 그대로입니다. 다만 그 표식이 말해주는 범위는 생각보다 좁습니다.
Anthropic has published how Claude's watermark works. The document went up on August 14.
It differs from what was reported three days earlier, particularly on editing.
Where the mark lands and where it does not is the substance of this document.
It does not insert characters
Anthropic states it plainly in the summary: nothing is added to the text, and there are no hidden characters.
So where does the mark live? In the moment a word is chosen. The model writes one word at a time, picking from a list of candidates for the next one.
The document's example is “The weather today is cold and…”. 'Sweet' is unlikely to follow; 'cloudy' and 'grey' are both natural. Choices like that, where either option works, happen countless times in a piece of writing, and the watermark leaves a pattern across them.
A reader cannot notice it; only a holder of the key can detect it. It is a different thing from the watermark on a banknote. It was never meant to be seen.
So did Anthropic invent this? The method is not new. It is a variant of SynthID-Text, published by Google DeepMind in Nature in 2024, in a line going back to Scott Aaronson's 2022 proposal. No extra tokens are used, so it does not raise cost.
Where it marks and where it does not
The watermark attaches only to words Claude chose. Everything else follows from that one line.
If you hand Claude something you wrote and ask only for grammar and punctuation, the only places a mark can attach are those few corrections. Anthropic writes that this may be too little to detect.
Translation is the opposite. Every word of a translation was chosen by Claude.
Code splits too. After “2 + 2 =” there is no equally good answer other than '4'. Where there is no choice, no watermark attaches. Comments, where wording can vary, can carry one, and the effect on the code itself is negligible.
Factual sentences are sparse as well. After “Newton's most famous work is the Principia…” the next words are fixed. Where being wrong is not allowed, there is nothing to choose.
So can you just erase it? Removal is a matter of degree. Light editing will not erase it entirely; rewriting every word will. Anthropic adds that in that case it is debatable whether the text is still AI-made at all.
So what does it prove
The watermark answers one question: whether Claude was involved at some point.
Anthropic states that it cannot distinguish “Claude wrote this” from “Claude heavily edited this”. It does not confirm that a human wrote something, and it says nothing about other AI systems.
Nor does it track anyone. Neither the watermark nor the key contains information identifying a user, an organisation or a conversation.
The reason for doing it is the EU AI Act. Since August 2, AI providers serving the EU market must mark AI-generated content. Anthropic signed the EU's code of practice on transparency for AI-generated content in July 2026, alongside about 190 signatories.
The verification tool is not out yet. A detection API is coming, with implementation details still being worked out. Images and files work differently: png, jpg and svg carry a cryptographically signed note in file metadata under the open C2PA standard.
Not only Claude
Is this only about Claude? The number of companies marking output is growing. Google has applied SynthID to text generated by the Gemini app and web since May 2024, and OpenAI is working to extend provenance marking to text, Newsis reported.
Anthropic will apply it to models released from now on, and roll it into existing Claude models over several months.
The direction is not uniform, though. On the 14th Google made visible watermarks on generated images, video and music a user setting, while keeping SynthID inside the content and C2PA provenance metadata. xAI has not yet announced any plan to apply a machine-readable invisible watermark to text written by Grok. It marks images and video visibly, but for text it goes only as far as advising users to label output as made with Grok, ZDNet Korea reported.
Fewer visible marks, machine-readable signals kept. That is where the industry is converging.
What this blog corrected
One addition. A piece on this blog dated August 11, based on Korean reporting, said a mark could remain even if you only had Claude proofread or translate your own writing.
Against the official document, translation holds but editing does not. If you are only fixing grammar, there is almost nothing for the mark to attach to. That article has been corrected.
Marks are spreading. What they can tell you is narrower than it sounds.
Sources · Anthropic Blog · The Dong-A Ilbo · Newsis · ZDNet Korea