Here is a hypothetical case. A teammate pastes a paragraph of changelog notes into your PR description. They got the text from an AI chat that had summarized a third-party web page. The English reads fine. The diff looks clean. CI is green. Two weeks later a security report points at one of the bullet points and asks why your release notes contain a 41-character invisible string in the U+E0000 range — the Unicode tag block. Nobody on the team typed it. Nobody on the team can see it in the rendered Markdown. It rode along with the copy-paste, survived your editor, your linter, your reviewer’s eyes, and your static site generator, and ended up checksummed into the artifact you shipped to customers. The same week, an unrelated audit flags one of your error messages because someone pasted a translated string from a rich-text editor and a U+202E RIGHT-TO-LEFT OVERRIDE came along with it; the string now reorders the visible characters of every error code printed after it on a terminal.
Try the Invisible Character Detector →
These are not exotic problems. They are the daily output of working with AI assistants, multilingual content, and rich-text sources. ZeroTool’s detector takes any pasted text and tells you exactly which code points are invisible, which category they belong to, and what the cleaned string looks like — all of it in the browser, with nothing uploaded.
What counts as “invisible”
The Unicode standard intentionally defines characters that render with zero width or that exert side effects on rendering without producing a glyph. They exist for legitimate reasons — Arabic shaping, Devanagari ligatures, soft line-break hints, emoji ZWJ sequences, file encoding markers — and become a problem only when they cross a boundary the author did not intend, such as plain-text export, source code, or a database column expecting ASCII.
The detector groups invisible code points into five categories. Each category has its own attack surface and its own legitimate use:
| Category | Code points | Legitimate use | Risk when smuggled in |
|---|---|---|---|
| Zero-width | U+200B ZWSP, U+200C ZWNJ, U+200D ZWJ, U+2060 WJ, U+FEFF BOM/ZWNBSP, U+3164 Hangul Filler, U+115F / U+1160 / U+FFA0 Hangul fillers, U+180E MVS, U+2061–U+2064 invisible math | Soft wrapping; Arabic / Indic shaping; emoji ZWJ sequences; file BOM | Identifier collisions, watermarking, parser drift, fingerprinting |
| Bidirectional | U+200E LRM, U+200F RLM, U+202A–U+202E LRE/RLE/PDF/LRO/RLO, U+2066–U+2069 LRI/RLI/FSI/PDI | Mixed LTR/RTL paragraphs, Arabic/Hebrew/Persian text | Trojan-Source (CVE-2021-42574) — reorders source code visually |
| Tag characters | U+E0000–U+E007F | Originally for language tagging in plain text (Unicode now deprecates U+E0001 LANGUAGE TAG and strongly discourages tag-based language tags); used today in emoji subdivision flag sequences | Hidden ASCII text (“ASCII smuggling”): prompt-injection instructions that people cannot see but LLMs can read |
| Variation selectors | U+FE00–U+FE0F (VS1–VS16), U+E0100–U+E01EF (VS17–VS256) | Select glyph variant for emoji (text vs emoji presentation) and CJK ideograph variants | Breaks string comparison, inflates byte length |
| Formatting | U+00AD SOFT HYPHEN, U+034F COMBINING GRAPHEME JOINER, U+206A–U+206F deprecated format controls | Suggested hyphenation points, grapheme cluster control | Survives paste into plain text; breaks substring search and tokenization |
Together the five categories cover the 4,174 code points that Unicode 18.0 marks Default_Ignorable_Code_Point in DerivedCoreProperties.txt. A code point that is not invisible can still be malicious — homoglyph attacks substitute Cyrillic а (U+0430) for Latin a (U+0061) and look identical at most font sizes. That is a different problem (confusable detection, specified in UTS #39) and is not what this tool addresses. The detector is strictly about code points that do not produce a glyph, or that produce only side effects on the rendering of other characters.
Why they matter — three real-world threats
Trojan-Source (CVE-2021-42574)
In November 2021 Nicholas Boucher and Ross Anderson at Cambridge published Trojan Source, demonstrating that almost every compiler, IDE, and code review tool at the time rendered bidirectional Unicode control characters according to the Unicode Bidirectional Algorithm — even inside comments and string literals in source code. By inserting RLI, LRI, PDI, and RLO controls, an attacker can author a source file in which the bytes the compiler sees say one thing, but the glyphs a reviewer sees say another.
A canonical example reorders a comment so that a return statement appears to be inside the /* ... */, while the compiler reads it as live code:
// JavaScript example, with U+202E (RLO) visualised as ⮜
const isAdmin = false;
/* Check if user is admin ⮜ begin admins only */
if (isAdmin) {
console.log("You are an admin.");
/* end admins only ⮜ */ }
Compilers added checks: Rust 1.56.1 introduced two deny-by-default lints, text_direction_codepoint_in_literal and text_direction_codepoint_in_comment (Rust security advisory). But such checks cover editors and compilers — not the documents, READMEs, Markdown files, configuration files, JSON blobs, or shell snippets that flow through the rest of your toolchain. A bidi-control hidden in a JSON config or a YAML release manifest is still invisible to most people reviewing it.
The fix on the detection side is mechanical: the entire bidi block is well-defined, and stripping it produces a string whose visible order is the same as its byte order. Run the tool over any inbound text that you cannot author yourself, and the bidi count tells you whether to look harder.
ASCII smuggling via tag characters
The Unicode tag block U+E0000–U+E007F was designed for language tagging in plain text (RFC 2482, 1999). The Unicode code chart for the block now marks U+E0001 LANGUAGE TAG as deprecated and says that using tag characters to convey language tags is strongly discouraged. The block survives in emoji subdivision flag sequences (the tag sequence used for 🏴 Scotland and similar). The other tag characters are invisible space: assigned code points that render to nothing, and do not interact with surrounding text.
That makes the block a near-perfect steganographic channel. Each ASCII byte can be encoded by adding U+E0000 to its code point, so U+E0020–U+E007E map 1:1 to the 95 printable ASCII characters. A 32-character token encodes into 32 invisible tag characters that ride inside any normal sentence.
In January 2024 Riley Goodside showed that instructions written in tag characters and hidden inside pasted text could prompt-inject ChatGPT. Johann Rehberger then published ASCII Smuggler, a tool that encodes and decodes these payloads, and showed that an LLM can also emit tag characters in its answers, so hidden text can flow out of a chat as well as into it. He describes the technique as a way to hide instructions for LLMs and to smuggle data in plain sight, and recommends that LLM applications filter tag characters out of prompts and responses.
Nothing public shows that tag characters are an AI watermark. OpenAI has not said that it marks ChatGPT output with tag characters. In April 2025 reports that o3 and o4-mini output contained unusual characters concerned U+202F NARROW NO-BREAK SPACE, not tag characters, and Rumi, the company that reported them, later wrote that OpenAI told it those characters are not a watermark but “a quirk of large-scale reinforcement learning”. No detector can tell from characters alone whether a text was written by AI.
If you paste web content, documents, or chat output into a prompt, a repository, or a publication, you want to know whether tag characters are present first. The detector flags the entire U+E0000–U+E007F range except inside the England, Scotland and Wales flags, labels each tag character with the ASCII letter it stands for, and removes the run with a single mode.
Copy-paste contamination
A frequent source is ordinary copy and paste. A Qiita post from August 2024 describes copying lines out of PowerPoint into VS Code: a U+200B sat before each line break, and after the file was saved as Shift_JIS every line ended in a question mark, because Shift_JIS has no zero-width space.
When that text leaves the rich-text environment and lands in a plain-text destination — a database varchar, a YAML file, a Markdown post, a code comment, an HTTP header — the invisible characters come along and quietly break things:
- Substring search misses matches:
"production"does not equal"pr\u200bduction". - Hash and signature checks fail intermittently because two visually identical strings produce different digests.
- Compiler errors point at the wrong column number because the source bytes are longer than the visible characters.
The lesson is general: anywhere text flows from a rich-text source into a security-relevant context, you need a way to see what is actually there.
How to detect and strip — the workflow
Open the Invisible Character Detector. Paste text into the input box. The line under it counts what was found, next to a Copy cleaned text button. Below are an annotated rendering that shows each invisible character as a labeled chip, a stats panel of categories and counts, and the strip mode selector.
Paste any text. The detection is synchronous and runs on every keystroke. You will see four useful pieces of information:
- Total count — how many invisible code points were found, broken down by category. A clean document reports zero.
- Per-character annotation — every invisible code point is shown inline as a chip with a short name such as
ZWSP,RLOorTAG-h. Hover a chip for its code point and full name. - Category counts — the stats panel counts each category, plus visible characters and total code points. A long run of tag characters in a row is probably hidden ASCII text; a single BOM at the start is usually a file encoding marker.
- Cleaned output — the same text with the selected category removed, ready to copy.
The strip mode selector has five positions:
- All — remove every invisible code point regardless of category. Use this when the source is plain text and there is no legitimate reason for any of these characters to be present. Most code, configuration files, JSON, YAML, and log lines fall in this bucket.
- Zero-width only — strip ZWSP, ZWNJ, ZWJ, WJ, BOM, Hangul filler, MVS, and invisible math. Preserve bidi controls (because RTL text may legitimately need them) and variation selectors (because emoji presentation depends on them). Use this when cleaning mixed-script writing where you want layout intent preserved.
- Bidi only — strip the bidirectional block exclusively. Use this for source code, configuration files, and anywhere the visible order must match the byte order, while keeping legitimate ZWJ sequences inside emoji or Devanagari intact.
- Tag only — strip the U+E0000–U+E007F range (the three flags above stay). Use this for pasted web or chat text where the only suspicious category is hidden tag text. Preserves everything else.
- Variation only — strip U+FE00–U+FE0F, U+E0100–U+E01EF and the Mongolian free variation selectors. Useful when emoji should fall back to text presentation, or when a string comparison fails because of an invisible selector.
Selecting a mode updates the cleaned output in place. Copy with the button, or download as .txt for binary-clean transport.
Detect and strip without the tool
The tool exists because clicking is faster than scripting. But the underlying detection is regex-trivial in any language. Below are three reference implementations you can drop into a CI step, a pre-commit hook, or a script that audits incoming user content.
The Python version uses only the standard library and prints a categorized count plus a cleaned string. Run it as python detect_invisible.py < input.txt:
import re
import sys
import unicodedata
CATEGORIES = {
"zero-width": r"[\u200B-\u200D\u2060-\u2064\uFEFF\u180E\u3164]",
"bidi": r"[\u200E\u200F\u202A-\u202E\u2066-\u2069]",
"tag": r"[\U000E0000-\U000E007F]",
"variation": r"[\uFE00-\uFE0F\U000E0100-\U000E01EF]",
"formatting": r"[\u00AD\u034F\u115F\u1160]",
}
def scan(text: str) -> dict[str, list[tuple[int, str, str]]]:
findings: dict[str, list[tuple[int, str, str]]] = {k: [] for k in CATEGORIES}
for name, pattern in CATEGORIES.items():
for match in re.finditer(pattern, text):
cp = match.group(0)
findings[name].append((
match.start(),
f"U+{ord(cp):04X}",
unicodedata.name(cp, "<unknown>"),
))
return findings
def strip_all(text: str) -> str:
combined = "|".join(p.strip("[]") for p in CATEGORIES.values())
return re.sub(f"[{combined}]", "", text)
if __name__ == "__main__":
src = sys.stdin.read()
report = scan(src)
total = sum(len(v) for v in report.values())
print(f"invisible code points: {total}")
for cat, hits in report.items():
if hits:
print(f" {cat}: {len(hits)}")
for offset, cp, name in hits[:5]:
print(f" @{offset} {cp} {name}")
sys.stdout.write(strip_all(src))
The JavaScript / TypeScript version targets Node 20+ and browsers. The same regexes work; the only twist is that JS source files need the u flag and surrogate-pair-aware syntax for code points above U+FFFF:
const CATEGORIES = {
"zero-width": /[\u200B-\u200D\u2060-\u2064\uFEFF\u180E\u3164]/gu,
"bidi": /[\u200E\u200F\u202A-\u202E\u2066-\u2069]/gu,
"tag": /[\u{E0000}-\u{E007F}]/gu,
"variation": /[\uFE00-\uFE0F\u{E0100}-\u{E01EF}]/gu,
"formatting": /[\u00AD\u034F\u115F\u1160]/gu,
};
const ALL = new RegExp(
Object.values(CATEGORIES).map(r => r.source).join("|"),
"gu"
);
export function detectInvisible(text) {
const findings = {};
for (const [name, re] of Object.entries(CATEGORIES)) {
findings[name] = [...text.matchAll(re)].map(m => ({
offset: m.index,
codePoint: "U+" + m[0].codePointAt(0).toString(16).toUpperCase().padStart(4, "0"),
}));
}
return findings;
}
export function stripInvisible(text) {
return text.replace(ALL, "");
}
If you want a one-line guard inside Bash — to fail a CI step on any tag character in a Markdown post, for example — GNU grep built with PCRE support works in a UTF-8 locale (on macOS, install it with Homebrew):
# Fail if any tag character (U+E0000–U+E007F) appears
if grep -P '[\x{E0000}-\x{E007F}]' "$file" >/dev/null; then
echo "tag characters detected in $file" >&2
exit 1
fi
# Strip every category in place with perl
perl -CSDA -i -pe '
s/[\x{200B}-\x{200D}\x{2060}-\x{2064}\x{FEFF}\x{180E}\x{3164}]//g;
s/[\x{200E}\x{200F}\x{202A}-\x{202E}\x{2066}-\x{2069}]//g;
s/[\x{E0000}-\x{E007F}]//g;
s/[\x{FE00}-\x{FE0F}\x{E0100}-\x{E01EF}]//g;
s/[\x{00AD}\x{034F}\x{115F}\x{1160}]//g;
' "$file"
perl -CSDA enables UTF-8 on STDIN, STDOUT, and @ARGV, which is the portable way to keep Perl from mangling multibyte input on the command line. The same script runs inside Git pre-commit hooks, GitHub Actions, and Vercel build steps without additional dependencies.
Pitfalls
Five edges to keep in mind when running invisible-character cleanup at scale:
Emoji ZWJ sequences are legitimate ZWJ. The family emoji 👨👩👧👦 is encoded as MAN U+200D WOMAN U+200D GIRL U+200D BOY — four base emoji glued together by three zero-width joiners. Stripping ZWJ from a string that contains emoji will turn that into four separate emoji rendered side by side. Same goes for 🏳️🌈 (white flag + ZWJ + rainbow) and the profession and hairstyle sequences. The detector flags ZWJ inside emoji because it has no way to distinguish “intentional sequence” from “smuggled byte” — visually, neither produces a glyph of its own. Use Bidi only or Tag only when cleaning text that contains emoji you want to preserve, or post-process by reapplying the canonical emoji sequences from a reference list.
File BOMs are sometimes intentional. Windows PowerShell 5.1 reads a script without a BOM in the legacy “ANSI” code page, so Microsoft recommends saving scripts that contain non-ASCII characters as UTF-8 with BOM. If your text came from a file rather than a clipboard, decide explicitly whether the leading BOM is meaningful before stripping it. The detector reports BOM as a zero-width code point regardless of position; you decide whether the report is a warning or an artifact.
Soft hyphens are normal in rich text. U+00AD is the recommended way to suggest hyphenation points to a rendering engine. A typeset PDF or an EPUB book may contain hundreds of them legitimately. Strip soft hyphens only when the target is plain text — code, configuration, database fields, log lines. Inside a typeset document, removing them degrades line-breaking quality without any security benefit.
Tag characters are not always hidden text. The U+E0000–U+E007F range still has one official use: emoji subdivision flag sequences. The Welsh flag 🏴 is composed of the black flag (U+1F3F4), the tag-encoded ISO subdivision code gbwls, and a CANCEL TAG (U+E007F) terminator. Stripping the entire tag block deletes those flags. A tag run between 🏴 and U+E007F is not proof of a flag, because anyone can put arbitrary tag text there. Only three sequences are recommended for general interchange (RGI_Emoji_Tag_Sequence in emoji-sequences.txt, Emoji 18.0): England gbeng, Scotland gbsct and Wales gbwls. The detector keeps exactly these three in every mode and flags every other tag character, including a fake flag that wraps hidden text.
Client-side cleanup does not fix upstream. If a CMS, a translation memory, or an LLM API is the source of the invisible characters, stripping them in the browser only cleans the copy you happen to be holding. The next copy from the same source has the same problem. Treat the detector as a microscope, not a filter — use it to confirm a hypothesis about the source, then put the actual strip step at the boundary you control (a webhook, a CI step, a pre-commit hook, a server-side normalisation routine using one of the implementations above).
A sixth pitfall worth mentioning: byte length is not character length is not visible width. A string with 50 visible characters and 80 invisible BMP characters such as U+200B has a String.length of 130 in JavaScript, a len() of 130 in Python, but a wcswidth of 50 in a terminal. Hash functions, content-length headers, database VARCHAR(N) limits, and authentication signatures all count the invisible ones too. If you compare strings normalised by visible width but stored by byte count, you will get false equality on inputs that should be distinct, or false inequality on inputs a human would call identical. NFC / NFKC normalisation in Unicode Normalization handles some cases (combining marks, compatibility decomposition) but does not remove invisible code points; stripping is a separate pass.
Comparison with other detectors
On 2026-09-29 we pasted the tool’s Hidden tag text sample, a chat reply with a family emoji and 25 tag characters, into two other tools.
Invisible Character Viewer labeled both zero-width joiners and every tag letter. Its Strip Invisible Characters button returned Thanks for the fix 👍 Team: 👨👩👧 See you Monday.: the hidden text was gone, and so were the joiners, so the family emoji split into three people. It also marks visible spaces such as U+00A0 and U+3000, which this detector leaves alone.
ASCII Smuggler decoded the tags into one readable sentence and counted 25 Unicode tags and 2 other invisible characters. It shows the decoded text for reading and does not produce a cleaned copy.
This detector removes one category at a time. With Tag only, the output keeps 👨👩👧 and drops the hidden sentence.
Further reading
Internal:
- Unicode Text Converter — turns text into styled Unicode letters such as 𝐁𝐨𝐥𝐝. Those are visible characters, so this detector does not flag them.
- String Escape — escapes and unescapes JavaScript, JSON and HTML entity strings. Its JavaScript mode writes U+200B as
\u200b, which is handy for putting a character you found into a test.
External:
- Trojan Source: Invisible Vulnerabilities — the Boucher & Anderson paper describing CVE-2021-42574 and CVE-2021-42694, with attack templates and mitigation guidance.
- Unicode Standard Annex #9 — Unicode Bidirectional Algorithm — the canonical specification for the bidi controls, including the embedding and isolate operators that Trojan-Source exploits.
- Unicode Technical Report 36 — Unicode Security Considerations — the standard’s own threat model for invisible characters, confusables, and identifier spoofing.
- Johann Rehberger — ASCII Smuggler Tool: Crafting Invisible Text and Decoding Hidden Codes — how tag characters hide instructions for LLMs, with an encoder/decoder and a demo of an LLM that emits hidden text.
- Unicode code chart: Tags (U+E0000–U+E007F) — the block’s character list, including the deprecation note on U+E0001 LANGUAGE TAG.
- RFC 2482 — Language Tagging in Unicode Plain Text — the original proposal for using tag characters as language tags.