What this does
Shows what is really inside a piece of text. Beyond the usual counts, it tells you how many characters, code points, and user-perceived characters (graphemes) there are, how many bytes the text takes in UTF-8, UTF-16, and UTF-32, and where any invisible characters hide.
Why the numbers differ
- Characters is the JavaScript
length: UTF-16 code units. An emoji like π counts as 2. - Code points are Unicode scalar values; π is one code point.
- Graphemes are what a reader sees as one character, counted with
Intl.Segmenter. A family emoji made of three emoji joined together is 1 grapheme, 5 code points, and 8 UTF-16 units. - Bytes depend on the encoding: ASCII letters take 1 byte in UTF-8, accented letters 2, most other scripts 3, and emoji 4. UTF-16 uses 2 or 4 bytes per code point.
Views
- Summary: counts plus encoding facts such as ASCII-only and Unicode normalization.
- Invisible: zero-width spaces, bidirectional controls, non-breaking and other unusual spaces, control characters, byte order marks, and Unicode tag characters, listed with line and column. One click removes the hidden ones.
- Code points: each character with its code point, UTF-8 bytes, and UTF-16 units.
- Whitespace: spaces, tabs, and line breaks drawn as visible symbols.
Notes
Browsers convert pasted line breaks to LF, so CRLF line endings are only detected in a file you open with βOpen fileβ. Large text is analysed in a background thread.
Privacy
This runs entirely in your browser. Nothing is uploaded.