TextWash

This page finds typographic ligatures (U+FB00–U+FB06: ff fi fl ffi ffl ſt) and unfolds them into plain letters via the "ligatures & fullwidth" rule. The sample below came out of a PDF. Everything runs in your browser — nothing is uploaded.

Input

X-ray hover a highlight for details

Cleaned output copied ✓

Cleaning rules

What are ligature characters?

Single Unicode characters that fuse letter pairs a typesetter would join: fi (U+FB01) for f+i, fl (U+FB02) for f+l, ffi (U+FB03) for f+f+i, and friends. In a rendered document they're pure typography. In extracted text, "finance" is a five-character string that will never match "finance".

Where they come from

Copying or extracting text from PDFs is the overwhelming source — the PDF stores the ligature glyph the typesetter used, and naive extraction hands it to you as-is. LaTeX output, InDesign exports and some ebook formats do the same.

What they break

Search, first and foremost: Ctrl+F, grep, database LIKE, and log search all miss words containing them. Résumé parsers and ATS systems are a notorious casualty — a PDF CV full of "office" and "certified" fails keyword matching invisibly. Spellcheckers flag correct words; deduplication and string equality fail; and code copied from a typeset document won't compile.

Fix it

Paste the text above. Ligatures are flagged as "other non-ASCII" in the x-ray, and the "Unfold ligatures & fullwidth chars" rule (NFKC normalization) expands each one to its plain letters — fi→fi, ffi→ffi — so search and matching work again. In your own pipeline: unicodedata.normalize("NFKC", text) in Python or text.normalize("NFKC") in JavaScript does the same.