>_devtools

UTF-8 / UTF-16 Encoder and Decoder

See the bytes of any text in UTF-8, UTF-16 LE and UTF-16 BE side by side, and decode hex bytes back to text.

What this does

Shows how a piece of text is stored as bytes in UTF-8, UTF-16 little endian and UTF-16 big endian, side by side, and decodes hex bytes in any of those encodings back into text. It is the quickest way to answer "why is this string 7 bytes?" or to read a UTF-16 dump from a Windows API, a database column or a network capture.

UTF-8 vs UTF-16

UTF-8 uses 1 byte for ASCII, 2 for most Latin, Greek and Cyrillic letters, 3 for most CJK characters and 4 for emoji. UTF-16 uses 2 bytes for everything in the Basic Multilingual Plane and 4 bytes (a surrogate pair) for characters above U+FFFF such as 🚀. UTF-16 comes in two byte orders: little endian stores é as e9 00, big endian as00 e9.

Byte order mark

Tick "Byte order mark" to prefix the output with the BOM (ef bb bf for UTF-8,ff fe or fe ff for UTF-16). When decoding, a BOM matching the selected encoding is removed. Decoding rejects invalid UTF-8 sequences, odd-length UTF-16 input and unpaired surrogates, naming the byte offset of the problem.

Related

To see the code points, categories and escapes of each character, use the Unicode inspector. For binary strings or hex with custom separators in other encodings use the text-to-hex and text-to-binary converters.

Privacy

The conversion happens in your browser. Nothing is uploaded or stored.