What Are Zero-Width Characters and Why Are They Hiding in Your AI Text?
Ever copied text from ChatGPT and noticed something felt... off? Maybe a headline wouldn't wrap correctly. Maybe a search-and-replace missed a word that was clearly there. Maybe your code editor threw a tiny fit for no obvious reason. I've seen this happen more than once, and the culprit often turns out to be something you can't even see: zero-width characters.
These invisible Unicode characters don't look like letters, spaces, or punctuation. They just sit there quietly inside text. And yes, they can show up in AI-generated writing from tools like ChatGPT, Claude, and Gemini. Sometimes they serve harmless formatting purposes. Sometimes they look a lot more like hidden markers.
So here's the deal: if you work with AI text, publish content, clean up copy, or move text between apps all day, it helps to know what these characters are and why they keep hitchhiking into your documents.
What are zero-width characters?
Zero-width characters are invisible Unicode characters. They exist in the text stream, but they don't display as normal visible symbols. Think of them as ghost characters. They count as data, not as visible ink.
The most common ones you'll run into are:
U+200B— Zero-Width Space or ZWSPU+200C— Zero-Width Non-Joiner or ZWNJU+200D— Zero-Width Joiner or ZWJU+FEFF— historically used as a BOM, sometimes called a BOM watermark when used as a hidden marker
U+200B is the one people mention most when talking about zero-width space AI issues. It acts like a possible break point in text without showing a visible space. That can be useful in typography. It can also be abused to split words invisibly.
ZWNJ and ZWJ matter in scripts where letter joining changes meaning or appearance, but they also show up in copied text where you really didn't expect them. And U+FEFF can appear at the beginning of a file as a byte order mark, though some systems also use it as a hidden separator inside text. That's where the phrase BOM watermark starts popping up.
A concrete example helps. These two words can look identical on screen:
watermark
watermark
The second version contains a U+200B in the middle. Visually, no difference. Under the hood, they are not the same string. That's the kind of thing that breaks exact matching, indexing, validation, and sometimes your patience.
Why AI companies use them
Now, here's where it gets interesting. Why would an AI company insert hidden characters AI text users can't even see?
There's a reason, and it usually comes down to economics, tracking, or legal risk.
First, tracking. If a provider wants to know whether text came from its model, invisible markers can act like breadcrumbs. Not a giant tattoo, more like a tiny signature. If enough patterns survive copy-paste, the company may be able to say, "Yep, this likely originated here."
Second, policy enforcement and abuse detection. AI companies spend serious money dealing with spam, scraping, republishing, and automated content laundering. If they can tag output in subtle ways, they gain one more signal for moderation systems.
Third, legal and commercial reasons. If a platform gets dragged into disputes over attribution, misuse, or content provenance, watermark-like signals become useful. Frankly, companies like having leverage. Even if a hidden marker isn't perfect evidence, it can still support internal analytics or product decisions.
Do all invisible Unicode characters in AI output mean deliberate watermarking? No. Some are legitimate formatting artifacts, especially when text passes through rich editors, browsers, and multilingual systems. But some patterns feel a little too tidy to be accidental. If you've ever found repeated U+200B characters in suspiciously regular places, you know what I mean.
How these characters travel when you copy and paste
Here's the annoying part: zero-width characters move with the text. You copy a paragraph from an AI chat window, paste it into Google Docs, then into WordPress, then into Slack, and the invisible bits often come along for the ride like freeloaders.
Because they're valid Unicode, most apps preserve them. They aren't "broken" characters in the usual sense. They're perfectly legal code points. That means your browser, CMS, editor, email client, and even programming language may keep them intact unless something explicitly strips them out.
You might notice this when:
- search doesn't find a word you can clearly see,
- a URL or slug behaves strangely,
- text comparisons fail,
- tokenizers split words oddly,
- or a script counts characters differently than expected.
I've also seen pasted AI text look fine in one app and then break in another. That's classic invisible character behavior. Very polite on the surface, mildly chaotic underneath.
How to detect zero-width characters
You don't need to be a Unicode nerd to spot them, though a little curiosity helps.
1. Use an inspector or online paste tool
Paste suspicious text into a Unicode viewer or character inspector. Some tools highlight non-printing characters directly. If you want a quick purpose-built option, aiwatermarksremover.com is handy for checking and cleaning text copied from AI systems.
2. Check the hex or code points
In a hex editor or developer-friendly text tool, those invisible marks will appear as actual Unicode values. For example, U+200B may show up in UTF-8 as E2 80 8B.
3. Use a small script
If you're comfortable with scripting, this is the fastest way. For example, in JavaScript:
const text = "watermarktestdemo";
const matches = text.match(/[]/g);
console.log(matches);
And in Python:
text = "watermarktestdemo"
hidden = [hex(ord(ch)) for ch in text if ch in ""]
print(hidden)
If either script prints values, you've got invisible passengers.
How to remove them
Removing zero-width characters is usually simple once you know they're there.
Manual and tool-based cleanup
- Paste into a cleanup tool that strips non-printing Unicode
- Use your code editor's regex replace
- Paste as plain text when possible
- Normalize content before publishing or processing
A basic regex does the trick in many editors:
[]
Replace that with nothing, and you'll remove ZWSP, ZWNJ, ZWJ, and the common BOM watermark character.
If you deal with AI output often, it's worth building this into your workflow. A lightweight text sanitizer, a CMS plugin, or even a clipboard-cleaning shortcut can save you from weird formatting bugs later. Another practical option is www.aiwatermarksremover.com, especially if you don't want to mess with regex every time.
A practical example with actual Unicode points
Let's say you copy this sentence from an AI chatbot:
The best results come from clean input and careful editing.
Looks harmless. But the hidden version might actually be:
The best results come from clean input and careful editing.
That includes:
U+200Bafter "The"U+200Cafter "come"U+200Dafter "input"U+FEFFbefore the period
To your eyes, it's normal text. To a parser, it's a different string than the one you thought you copied. That matters if you compare versions, hash content, build search indexes, or process prompts programmatically.
Does this really matter?
My honest take: for casual use, maybe not much. If you're pasting a draft into a note app and moving on with your life, invisible Unicode characters probably won't ruin your afternoon.
But if you publish content, train systems, process text in bulk, analyze AI output, or care about clean data, yes, it matters. A lot more than people think. Tiny hidden marks can create surprisingly dumb bugs. And the older I get, the less patience I have for bugs caused by characters no human can even see.
So no, you don't need to panic every time you copy text from ChatGPT, Claude, or Gemini. Just know that these markers exist. Check for U+200B, ZWNJ, ZWJ, and stray U+FEFF when something feels weird. If text behaves strangely, trust that instinct. It's often not you. It's the ghosts in the Unicode.