A form generator stopped producing documents. Not slowly, and not for everyone — four records out of thousands returned a generic "could not fill that form" and the rest were fine.
The cause was two letters in somebody's first name.
Why one character loses the whole document
pdf-lib's StandardFonts write WinAnsi, an 8-bit encoding covering roughly Latin-1 plus a couple of dozen typographic extras. When it meets a character outside that set, it does not substitute a placeholder or drop it. It throws:
WinAnsi cannot encode "ff" (0xfb00)
That name was "Jeff". Except the two f's were not two f's — they were U+FB00, the single typographic ligature glyph, which is exactly what you get when someone pastes a name out of another PDF.
Here is the part that makes it expensive. The throw does not happen when you set the field. It happens inside doc.save(), while appearance streams are generated for every field in the document. So the failure is not "this field is wrong". It is "there is no document" — and the stack trace points at save(), nowhere near the record that caused it.
One character, in one field, loses everything.
Half the offenders are invisible
Across the four broken records, the culprits were:
- a black circle (U+25CF) someone had used as a bullet
- the ff ligature (U+FB00) from a pasted name
- a byte-order mark (U+FEFF)
- a zero-width space (U+200B)
Two of those render as nothing at all. You can open the record, read every field, copy the text into a document, and see absolutely nothing wrong — because there is nothing to see. That is why this belongs in a test rather than in a habit of checking carefully.
The obvious fix is a worse bug
The natural instinct is to normalize the string. NFKC turns ff into ff and fi into fi, which is precisely the problem you have. One line, done.
Do not do this.
NFKC is a compatibility normalization. It rewrites characters that were already perfectly encodable, and one of those rewrites is quietly destructive:
| Input | After blanket NFKC | WinAnsi could encode the original? |
|---|---|---|
| ½ | 12 | Yes |
| … | ... | Yes |
| ² | 2 | Yes |
| ™ | TM | Yes |
Look at the first row. ½ decomposes to 1⁄2 — and the fraction slash in the middle is itself non-WinAnsi, so it gets dropped by the same pass that was supposed to be helping. ½ becomes "12".
On a form describing a physical thing, a half-inch part silently becoming a twelve-inch one is a far worse outcome than the encoding error you started with. It does not throw. It does not log. It produces a valid, professional-looking document containing a wrong number.
The other three rows are less dangerous but equally pointless: WinAnsi carries all of them, so normalizing them is pure loss.
The rule that falls out of that
Apply NFKC only to characters that are already unencodable. Never as a blanket pass.
And normalize to NFC first, so that an "e" followed by a combining acute recomposes into é — which WinAnsi does carry — instead of surviving as a decomposed pair that later loses its accent.
The whole algorithm, in order:
function pdfSafe(text) {
let out = "";
for (const raw of text.normalize("NFC")) {
const ch = raw in FOLD ? FOLD[raw] : raw; // known shapes and invisibles
if (ch === "") continue;
if (encodable(ch.codePointAt(0))) { out += ch; continue; }
// Only now, and only on this character:
for (const n of ch.normalize("NFKC")) {
if (encodable(n.codePointAt(0))) { out += n; continue; }
// An accent we cannot carry is worth less than the letter under it.
for (const d of n.normalize("NFKD").replace(/\p{M}/gu, "")) {
if (encodable(d.codePointAt(0))) out += d;
}
}
}
return out;
}
The FOLD table handles what normalization will not: decorative shapes to a real bullet, prime and double prime to ' and " so that 42′ 8″ survives as feet and inches rather than becoming a bare number, the true minus sign to a hyphen, and the invisible characters to the empty string.
One subtlety in the encodable test
WinAnsi is not simply Latin-1. The byte range 0x80–0x9F, which is control codes in Unicode, is where WinAnsi keeps its typographic characters — curly quotes, en and em dashes, the ellipsis, the euro, the bullet, the trademark sign.
Those sit at unrelated Unicode codepoints, so they cannot be derived from a range. They have to be listed:
const WINANSI_HIGH = new Set([
0x20ac, 0x201a, 0x0192, 0x201e, 0x2026, 0x2020, 0x2021, 0x02c6, 0x2030,
0x0160, 0x2039, 0x0152, 0x017d, 0x2018, 0x2019, 0x201c, 0x201d, 0x2022,
0x2013, 0x2014, 0x02dc, 0x2122, 0x0161, 0x203a, 0x0153, 0x017e, 0x0178,
]);
Miss this and every curly quote and em dash in pasted text — which is to say, most pasted text — gets needlessly mangled.
Fix it at the boundary, not in the data
It is tempting to clean the source records instead. That does not hold. If the text arrives from systems where people paste out of Word, Outlook and other PDFs all day, new unencodable characters will keep arriving indefinitely, and a cleaning pass is a race you re-run forever.
Handle it where text meets the page, and route every call that draws variable text with a standard font through the same function — not only the one that broke. There were five such call sites in this codebase and only one had failed yet.
Write the test, because manual testing cannot see this
The first version of this fix shipped for one deploy cycle before the ½ problem was noticed. It looked correct. It was correct for every case anyone thought to try by hand, because the failure mode is a plausible number rather than an error.
Assertions worth having, all of which run offline with no database and no credentials:
ligatures become plain letters
pdfSafe("Jeff") === "Jeff"
pdfSafe("file flag") === "file flag"
invisible characters go quietly
pdfSafe("hidden") === "hidden"
pdfSafe("AB") === "AB"
shapes keep their meaning rather than vanishing
pdfSafe("● one") === "• one"
pdfSafe("42′ 8″") === "42' 8\""
what WinAnsi already carries is left alone
pdfSafe("½") === "½"
pdfSafe("café … ™") === "café … ™"
That last block is the one that matters. It is the difference between a fix and a quieter bug.
The general shape
Two things here generalize past pdf-lib.
First, an error thrown during a batch operation tells you nothing about which input caused it. The message named the character but not the field, the record or the reason it was there. If a library validates late, the stack trace will point at the library rather than at your data.
Second, the obvious fix for an encoding problem is usually too broad. Normalization is lossy by design, and "make everything safe" and "keep everything that was already correct" are different goals. Applying the aggressive transform only where the safe path has already failed costs a few more lines and is the difference between a document that fails loudly and one that lies.