Skip to content
augur

What it finds

A Content Credential hidden in text

The C2PA specification defines a way to carry a manifest in plain text: the whole manifest encoded as Unicode variation selectors, rendering as nothing, sitting at the end of the writing. It survives a copy and a paste, which is exactly the point of putting it there.

C2PATXTvariation selectorsU+FE00text manifestinvisible AI watermark

A paragraph that knows where it came from

$ augur scan marked-post.txt
marked-post.txt (text): 2 finding(s)

INVISIBLE
   [concern] offset 182 — 1473 invisible character(s): ZERO WIDTH NO-BREAK SPACE (BOM)
       in context: this year. The rest of this note is background.⏎⟨BOM⟩⟨VS68⟩⟨VS51⟩⟨VS81⟩⟨VS66⟩⟨VS85⟩⟨VS89⟩⟨VS85⟩⟨VS1⟩…

PROVENANCE
   [notice] offset 182 — C2PA Content Credential in this text (variation selectors, 1459 bytes) — Lumen Assistant 4, generated by a trained model
       generated by=Lumen Assistant 4
       source=generated by a trained model
       binding=matches — the file still hashes to what was signed (sha256)
       actions=c2pa.created
       asset title=post.txt
       asset format=text/plain
       signed by=Lumen Signing (Lumen), certificate issued by Lumen Signing
       signature=ES256
       claim=version 1, 2 assertions
       assertions=c2pa.actions, c2pa.hash.data

Two findings about the same 1,473 characters, and the difference between them is the point. The first says a run of invisible characters is there. The second says what it is: a signed statement that a model wrote this paragraph, still matching the words it was signed over.

Not the same thing as a statistical watermark

A statistical watermark lives in which ordinary words a generator chose. Nothing is hidden in the characters, and only whoever holds the key can test for it — so any tool claiming to find one by reading your text, or to have removed one, is guessing.

This is the other kind of mark, and it is the readable kind: characters, in the file, saying what made it. That is why augur reports it rather than shrugging.

A word changed after signing

$ augur scan marked-post-edited.txt
marked-post-edited.txt (text): 3 finding(s)

INVISIBLE
   [concern] offset 182 — 1473 invisible character(s): ZERO WIDTH NO-BREAK SPACE (BOM)
       in context: this year. The rest of this note is background.⏎⟨BOM⟩⟨VS68⟩⟨VS51⟩⟨VS81⟩⟨VS66⟩⟨VS85⟩⟨VS89⟩⟨VS85⟩⟨VS1⟩…

PROVENANCE
 * [concern] offset 182 — the Content Credential no longer matches this text
       algorithm=sha256
       claim says=85742285cac5…
       text hashes to=63acde0f3b61…
   [notice] offset 182 — C2PA Content Credential in this text (variation selectors, 1459 bytes) — Lumen Assistant 4, generated by a trained model
       generated by=Lumen Assistant 4
       source=generated by a trained model
       binding=DOES NOT match — the file changed after it was signed (sha256)
       actions=c2pa.created
       asset title=post.txt
       asset format=text/plain
       signed by=Lumen Signing (Lumen), certificate issued by Lumen Signing
       signature=ES256
       claim=version 1, 2 assertions
       assertions=c2pa.actions, c2pa.hash.data

* not removable — reported and left in place

Excerpt: the provenance section of a three-finding scan. One word of the visible text was changed — April to March — and the credential stops matching. For a document meant to be edited that is ordinary; for one forwarded as what the model wrote, it is not.

Taking it out

$ augur clean marked-post.txt --categories=provenance
wrote marked-post.clean.txt
verified: 1 removed, 0 finding(s) deliberately left in place

The credential is characters in a text file, so removing it is deleting exactly those bytes: the visible writing is byte-identical afterwards. Whether to remove it is a decision — it is a record of where the text came from, and taking it out takes that with it.

$curl -fsSL https://raw.githubusercontent.com/dejo1307/augur/main/install.sh | sh

Related:C2PA Content CredentialsInvisible characters in AI textThe same encoding, used dishonestlyeverything it looks for