Text encodings explained — UTF-8, legacy code pages and the BOM
Published:
Computers don’t store letters — they store numbers (bytes). An encoding is the rule that maps characters to bytes. If a file is saved with one rule and opened with another, the text comes out garbled. That garbage has a name: mojibake.
The encodings you’ll meet
| Encoding | What it is | Bytes for “é” | Bytes for “한” |
|---|---|---|---|
| ASCII | The original 128 characters: English letters, digits, punctuation | — | — |
| Windows-1252 | The Western European “ANSI” code page on Windows | 1 | — |
| Other code pages | Language-specific legacy encodings: CP949 (Korean), Shift_JIS (Japanese), GB18030 (Chinese), Windows-1251 (Cyrillic) | varies | 2 in CP949 |
| UTF-8 | The universal Unicode encoding used by almost the entire web | 2 | 3 |
| UTF-16 | A Unicode encoding used inside Windows and some apps | 2 | 2 |
Plain English letters use the same bytes in all of these, which is why English text survives while accents and other scripts break.
What is a BOM?
A byte order mark (BOM) is a short marker at the very start of a file. For UTF-8 it is the three bytes EF BB BF, and it acts as a label saying “this file is UTF-8”.
- Pro: Excel and older Windows programs recognise UTF-8 files correctly when the BOM is present.
- Con: some programming tools and configuration parsers treat it as a stray invisible character at the start of the file.
The usual rule: add a BOM to CSV files meant for Excel, and leave it out for files read by software. Our text-output converters let you choose under Options.
Diagnosing garbled text
The shape of the garbage tells you what went wrong:
| What you see | What happened |
|---|---|
é, ü, ’ |
UTF-8 text read as Windows-1252 |
� repeated |
A legacy-encoded file read as UTF-8 (invalid byte sequences) |
Cyrillic or CJK symbols in the middle of English words, like M黮ler |
A Western file read as a multi-byte Asian encoding |
| The same garbage getting longer each time | Text that was mis-decoded and then saved again — double encoding |
The crucial rule: don’t save a file while it looks garbled. Saving writes the wrong characters permanently. Close the file and reopen it with the right encoding.
Why old text files are still a problem
Text files, subtitles and e-books from the 2000s were usually saved in the system’s legacy code page, because that was Notepad’s default. Modern apps and phones assume UTF-8, so these files often open as gibberish.
The TXT to EPUB, TXT to DOCX and CSV to XLSX converters detect the encoding in this order:
- If the file starts with a BOM, follow it.
- If the bytes are valid UTF-8, use UTF-8.
- Otherwise try CP949 (Korean), Shift_JIS (Japanese), GB18030 (Chinese) and Windows-1251 (Cyrillic), and pick the one whose result reads as natural text in that script.
- Fall back to Windows-1252 (Western European).
If detection ever guesses wrong, you can pick the encoding manually under Options.
How to check a file’s encoding
- Windows Notepad: the status bar in the bottom-right shows
UTF-8,UTF-8 with BOMorANSI(your system’s legacy code page). - VS Code: the encoding is shown in the bottom-right; click it to reopen the file with a different encoding.
Summary
Garbled text means the saving rule and the reading rule differ. Save new files as UTF-8, add a BOM to CSVs meant for Excel, and convert old legacy-encoded files once with an encoding-aware tool so they display correctly everywhere.