Skip to content
augur

What it finds

Homoglyphs and lookalike letters

The same shape from a different alphabet. It reads correctly to you and matches nothing: not a search, not a filter, not an allowlist entry.

confusablesCyrillic а U+0430IDN homographfullwidth letters

Three kinds, three lines

$ augur scan allowlist.md
allowlist.md (text): 3 finding(s)

CONFUSABLE
 * [alarm] offset 80 — "раураӏ" is entirely Cyrillic — it reads as "paypal"
 * [concern] offset 20 — "pаssword" mixes Cyrillic and Latin
 * [concern] offset 122 — "𝐢𝐠𝐧𝐨𝐫𝐞" is mathematical bold letters — it reads as "ignore"

* not removable — reported and left in place

One word mixes alphabets. One is wholly Cyrillic, so nothing about it is mixed and every letter is a lookalike, which is why it is the alarm. One is Latin in an alternate Unicode copy of itself. All three are reported by what they read as, not by codepoint.

Why augur will not fix these for you

Every one carries a *: reported, never removed automatically. Which alphabet a word was meant to be written in is a judgment only you can make. It might be an attack, and it might be somebody’s name spelled the way they spell it.

The place this bites hardest is a permission allowlist or a filter rule. A lookalike letter in one means the rule you believe you wrote silently never matches anything, and nothing anywhere reports an error.

The false positive you will meet

Greek beside Latin is an attack in a password and ordinary notation in a maths document. augur reports it and cannot yet tell the difference, so it tells you what it saw and leaves the call with you.

$curl -fsSL https://raw.githubusercontent.com/dejo1307/augur/main/install.sh | sh

Related:Scanning a repositoryTrojan Source and bidi overrideseverything it looks for