October 4, 2026 · Yunus Emre Vurgun
What Is UTF-8? Character Encodings Explained Simply
UTF-8 is a way of storing Unicode text as bytes: it uses one byte for plain ASCII, two bytes for most European and Middle Eastern scripts, three for most Asian scripts, and four for emoji and rare symbols. It is backward compatible with ASCII, it needs no byte-order mark, and it has won so completely that over 98 percent of the web uses it. If you only remember one rule, make it this: save, serve, and store text as UTF-8 unless you have a specific reason not to.
Bytes vs characters: the core idea
Computers store bytes; humans read characters. An encoding is the agreed mapping between the two, and every encoding headache comes from two programs disagreeing about which mapping applies. ASCII, the 1960s ancestor, defined 128 characters in 7 bits — English letters, digits, punctuation, and control codes — and every byte with the top bit set was left undefined. That worked until anyone needed an accented letter, at which point each region invented its own "extended ASCII": Latin-1 for Western Europe, Shift_JIS for Japanese, KOI8-R for Russian, and dozens more. A file was only readable if you already knew which of these the author used.
Unicode solved the character side by assigning every character in every script a permanent number called a code point — U+0041 for "A", U+00E9 for "é", U+1F600 for the grinning emoji. But a code point is still not bytes: UTF-8, UTF-16, and UTF-32 are three different encodings of the same code points, i.e. three ways to serialize those numbers. Confusing "Unicode" (the catalog) with "UTF-8" (one serialization) causes half of all encoding misunderstandings, so keep the two layers separate in your head. The companion guide to units, dates, and encodings puts this in the wider context of representing real-world values in data.
How UTF-8 encodes a character
UTF-8 is variable-width: common characters get short byte sequences and rare ones get long ones. The first byte announces how many bytes follow, and every continuation byte starts with the bits 10, which is what makes the format self-synchronizing — you can always find character boundaries by scanning bytes.
| Code point range | Byte 1 | Bytes used | Covers |
|---|---|---|---|
| U+0000 – U+007F | 0xxxxxxx | 1 byte | ASCII: English, digits, punctuation |
| U+0080 – U+07FF | 110xxxxx | 2 bytes | Latin extensions, Greek, Cyrillic, Arabic, Hebrew |
| U+0800 – U+FFFF | 1110xxxx | 3 bytes | Chinese, Japanese, Korean, most symbols |
| U+10000 – U+10FFFF | 11110xxx | 4 bytes | Emoji, rare scripts, historic characters |
Two examples make it concrete. The letter "é" is U+00E9, which falls in the two-byte range and encodes as the bytes C3 A9. The grinning emoji U+1F600 needs four bytes: F0 9F 98 80. You can verify this yourself in seconds:
$ python3 -c "print('é'.encode('utf-8'), '😀'.encode('utf-8'))"
b'\xc3\xa9' b'\xf0\x9f\x98\x80'Notice what backward compatibility means here: every ASCII file is already valid UTF-8, byte for byte. That single property let the whole world migrate without converting existing data, and no competing Unicode encoding can claim it.
Why UTF-8 beat UTF-16, UTF-32, and Latin-1
Each rival encoding has a real drawback that UTF-8 avoids. Latin-1 and the other legacy code pages simply cannot represent most of the world's scripts — one byte caps you at 256 characters. UTF-32 uses four bytes for everything, quadrupling the size of English text for zero benefit. UTF-16 uses two or four bytes, which is compact for Asian text but wastes space on ASCII-heavy content like HTML tags, JSON keys, and source code — and it suffers byte-order ambiguity, requiring a byte-order mark or out-of-band agreement about endianness.
| Encoding | Bytes per char | ASCII compatible? | Main drawback |
|---|---|---|---|
| UTF-8 | 1–4 | Yes | 3 bytes for CJK text (vs 2 in UTF-16) |
| UTF-16 | 2 or 4 | No | Byte-order issues, ASCII wastes a byte |
| UTF-32 | Always 4 | No | Huge for mostly-ASCII data |
| Latin-1 | Always 1 | Yes (first 128) | Only 256 characters total |
UTF-16 survives mainly where history locked it in: JavaScript strings, Java and C# internals, and the Windows API all use it. That is worth knowing when counting "string length" — an emoji may count as two UTF-16 code units — but for files, APIs, and web pages, UTF-8 is the default the industry converged on.
Mojibake: what goes wrong and how to fix it
Mojibake — garbled text like "é" where "é" should be — happens when bytes written in one encoding are read as another. The classic case is UTF-8 bytes decoded as Latin-1: the two bytes C3 A9 display as the two Latin-1 characters "Ã" and "©". Diagnosis is a matter of recognizing the signature:
| You see | Probably means | Fix |
|---|---|---|
| é, ü, à | UTF-8 bytes read as Latin-1 | Re-decode the original bytes as UTF-8 |
| Question marks or � | Decoder replaced unknown bytes | Data may be lost; fix the pipeline, re-ingest |
| 䏿–‡ | UTF-8 CJK bytes read as Latin-1 | Re-decode as UTF-8 |
| Every non-ASCII char breaks | An ASCII-only step in the chain | Find the step (old driver, misconfigured server) and set UTF-8 |
Prevention beats repair. Declare UTF-8 at every layer: a <meta charset="utf-8"> tag in HTML, charset=utf-8 on HTTP Content-Type headers (see the content types guide), encoding="utf-8" when opening files in Python, and utf8mb4 — not the broken 3-byte utf8 — in MySQL. The file command guesses encodings decently (file -i suspect.txt), and iconv -f latin1 -t utf-8 converts legacy files when re-reading the original is impossible.
FAQ: is UTF-8 the same as Unicode?
Is UTF-8 the same as Unicode? No. Unicode is the catalog that assigns numbers (code points) to characters; UTF-8 is one encoding that turns those numbers into bytes. Saying "convert to Unicode" is meaningless until you name an encoding — you almost always mean "convert to UTF-8."
Should I use UTF-8 for everything? For text interchange — files, web pages, JSON APIs, database storage — yes. Every JSON API on this site, documented in the API reference, serves UTF-8 JSON. The exceptions are narrow: UTF-16 inside runtimes that require it, and legacy systems you cannot change, where you convert at the boundary.
Do I need a byte-order mark? No, and adding one often breaks things: Unix tools, JSON parsers, and many web servers treat a BOM as garbage. UTF-8 has no byte-order ambiguity, so the BOM buys nothing. Save as "UTF-8 without BOM."