What this does
Removes accents and other diacritical marks from letters, so Crème brûlée becomesCreme brulee. Useful for search keys, file names, comparisons, and systems that only accept plain ASCII letters.
How it works
The text is decomposed into base letters plus combining marks (Unicode NFD), the combining marks from the Combining Diacritical Marks blocks are deleted, and the result is recomposed (NFC). This handles accents, cedillas, umlauts, tildes, and carons on Latin, Greek, and Cyrillic letters, and works equally well on text that is already stored in decomposed form.
Letters that do not decompose
Some letters are not a base letter plus a mark, so stripping marks leaves them unchanged. With “Also transliterate” on, these are replaced: ß to ss, ẞ to SS, æ to ae, œ to oe, ø to o, đ to d, ð to d, þ to th, ł to l, ħ to h, and dotless ı to i. Turn it off to leave them as they are.
What it leaves alone
Scripts where marks are part of the letters, such as Hindi, Thai, and Arabic, are not changed, and neither are Chinese, Japanese, and Korean. This tool does not translate or romanize between scripts.
Privacy
This runs entirely in your browser. Nothing is uploaded.