>_devtools

Unicode Character Inspector

Inspect every code point of a string: U+ value, category, script, UTF-8 bytes, UTF-16 units, and invisible characters.

What this does

Breaks text into Unicode code points and shows, for each one, the character, itsU+ code point, general category, script, the bytes it occupies in UTF-8, and the 16-bit code units it occupies in UTF-16. Totals at the top count code points, grapheme clusters (what a reader sees as one character), UTF-16 code units (what JavaScript's.length reports) and UTF-8 bytes (what gets sent over the wire).

Why the numbers differ

An emoji such as 👋🏽 is two code points (the waving hand and a skin-tone modifier), four UTF-16 code units and eight UTF-8 bytes, yet it displays as one grapheme. An "é" can be one code point (U+00E9) or two (e followed by a combining accent U+0301); both look identical. Invisible characters such as the zero-width space (U+200B), no-break space (U+00A0) and right-to-left marks are highlighted so you can find the stray character that breaks a comparison, a regular expression or a JSON key.

Names, categories and scripts

Categories (Lu, Nd, Sc, …) and scripts come from the Unicode data built into your browser, so they track the Unicode version your browser supports. Character names are included for ASCII, common spaces, dashes, quotes, bidi controls and a few frequent symbols; other characters show a dash instead of a name because a full name table would add several hundred kilobytes. Switch on "Show escapes" to get the JavaScript, HTML, CSS and URL-encoded form of each character.

Limits

Up to 200,000 characters are analyzed and the first 2,000 code points are listed in the table; the totals always cover everything analyzed. Lone surrogates (broken UTF-16) are reported as category Cs and cannot be encoded as UTF-8.

Privacy

The analysis runs in your browser. Nothing you paste is uploaded or stored.