What this does
Shows how a piece of text is stored as bytes in UTF-8, UTF-16 little endian and UTF-16 big endian, side by side, and decodes hex bytes in any of those encodings back into text. It is the quickest way to answer "why is this string 7 bytes?" or to read a UTF-16 dump from a Windows API, a database column or a network capture.
UTF-8 vs UTF-16
UTF-8 uses 1 byte for ASCII, 2 for most Latin, Greek and Cyrillic letters, 3 for most CJK characters and 4 for emoji. UTF-16 uses 2 bytes for everything in the Basic Multilingual Plane and 4 bytes (a surrogate pair) for characters above U+FFFF such as 🚀. UTF-16 comes in two byte orders: little endian stores é as e9 00, big endian as00 e9.
Byte order mark
Tick "Byte order mark" to prefix the output with the BOM (ef bb bf for UTF-8,ff fe or fe ff for UTF-16). When decoding, a BOM matching the selected encoding is removed. Decoding rejects invalid UTF-8 sequences, odd-length UTF-16 input and unpaired surrogates, naming the byte offset of the problem.
Related
To see the code points, categories and escapes of each character, use the Unicode inspector. For binary strings or hex with custom separators in other encodings use the text-to-hex and text-to-binary converters.
Privacy
The conversion happens in your browser. Nothing is uploaded or stored.