TextWash

This page cleans data copied from spreadsheets and CSVs: non-breaking spaces, narrow spaces inside numbers, Unicode minus signs and directional marks are flagged and normalized. The sample below fails float parsing twice. Everything runs in your browser — nothing is uploaded.

Input

X-ray hover a highlight for details

Cleaned output copied ✓

Cleaning rules

The value is right there. The lookup says it isn't.

Data that travels through Excel, Google Sheets, web pages or Windows dialogs picks up characters that look like spaces and dashes but aren't: non-breaking spaces (every   on a copied web table), narrow no-break spaces as thousands separators (42 000 — standard French formatting, and what ICU emits), the typographic minus sign U+2212 on negative numbers from Wikipedia and calculators, left-to-right marks inside dates copied from Windows, and the full-width space U+3000 from CJK input methods.

What it breaks

VLOOKUP and MATCH can't find values padded with a non-breaking space. float("−42.5") raises ValueError; parseFloat returns NaN; SQL imports store the number as text and sums silently skip it. TRIM() doesn't remove U+00A0, so "cleaning" the column changes nothing. Dates with embedded directional marks fail strptime with "unconverted data remains" even though they print perfectly.

Fix it

Paste the offending rows (or drop the whole CSV file — up to 5 MB, processed locally, never uploaded). Every exotic space, dash and invisible mark is flagged in the x-ray and normalized in the output; download saves it as name.clean.csv with line endings preserved. For a recurring pipeline, normalize at ingest: replace \u00a0, \u202f and \u2212 before parsing numbers.