Skip to content
augur

What it finds

Per-copy watermarks in a document

Send eleven people a confidential document and you can tell which of them leaked it, without changing a visible word. One invisible character a line, varied per recipient, is enough — and every individual character is deniable.

canary trapleak tracingper-recipient watermarktraitor tracing

The finding is the pattern, not the character

$ augur scan --min-severity=alarm briefing.txt
briefing.txt (text): 1 finding(s)

FINGERPRINT
 * [alarm] offset 84 — 2 kinds of zero-width character (U+200B, U+200C) appear 16 times across 16 lines, about one per line
       alphabet=U+200B ZERO WIDTH SPACE ×8, U+200C ZERO WIDTH NON-JOINER ×8
       occurrences=16
       lines=16
       cadence=about one per line
       offsets=84, 132, 196, 242, 303, 364, 418, 473, … and 8 more

* not removable — reported and left in place

Sixteen lines, one of two invisible characters after the first word of each. That is sixteen bits, which distinguishes sixty-five thousand copies. At the default severity the same file reports eighteen findings, sixteen of which are single characters worth nothing on their own.

Why a list of characters is not enough

Every other detector asks what a character is. That question has no useful answer for one zero-width space: it is a paste artefact, it is nothing, it is a notice at most, and calling it more would teach you to ignore the tool.

Two hundred of them, one to a line, through a document with no other reason to contain any, is not two hundred artefacts. It is one mark. The property that makes it a mark — sparse, regular, spread across the document — belongs to the set and to no member of it, and a findings list sorted by position is the one view that cannot show it.

What to do once you know

augur clean removes the characters, and the mark leaves with them: the fingerprint finding is reported as not removable because it describes a pattern rather than a thing in the file, and there is nothing left to describe once the characters are gone.

Whether removing it is wise is a separate question with an answer that depends entirely on your situation, and this tool has no view on it. Knowing the mark is there is the part that was hard.

What this does not catch

Marks carried in wording rather than in characters — a synonym swapped per recipient, a sentence reordered, a comma moved. Nothing in the file distinguishes those from writing, and no tool reading one copy can find them.

Statistical watermarks in generated text, which are a property of word choice across a whole passage. Watermarks in pixels rather than in text. A mark made of exotic spaces rather than zero-width characters is reported character by character but does not currently trip this detector, so read the whitespace notices with the same eye.

A clean report here means this detector found no distribution it recognises. It is not a statement that your copy is unmarked.

$curl -fsSL https://raw.githubusercontent.com/dejo1307/augur/main/install.sh | sh

Related:Zero-width charactersInvisible characters in AI texteverything it looks for