Skip to content
FreeConvertter

Text encodings explained — UTF-8, legacy code pages and the BOM

Published:

Computers don’t store letters — they store numbers (bytes). An encoding is the rule that maps characters to bytes. If a file is saved with one rule and opened with another, the text comes out garbled. That garbage has a name: mojibake.

The encodings you’ll meet

Encoding What it is Bytes for “é” Bytes for “한”
ASCII The original 128 characters: English letters, digits, punctuation — —
Windows-1252 The Western European “ANSI” code page on Windows 1 —
Other code pages Language-specific legacy encodings: CP949 (Korean), Shift_JIS (Japanese), GB18030 (Chinese), Windows-1251 (Cyrillic) varies 2 in CP949
UTF-8 The universal Unicode encoding used by almost the entire web 2 3
UTF-16 A Unicode encoding used inside Windows and some apps 2 2

Plain English letters use the same bytes in all of these, which is why English text survives while accents and other scripts break.

What is a BOM?

A byte order mark (BOM) is a short marker at the very start of a file. For UTF-8 it is the three bytes EF BB BF, and it acts as a label saying “this file is UTF-8”.

  • Pro: Excel and older Windows programs recognise UTF-8 files correctly when the BOM is present.
  • Con: some programming tools and configuration parsers treat it as a stray invisible character at the start of the file.

The usual rule: add a BOM to CSV files meant for Excel, and leave it out for files read by software. Our text-output converters let you choose under Options.

Diagnosing garbled text

The shape of the garbage tells you what went wrong:

What you see What happened
é, ü, ’ UTF-8 text read as Windows-1252
� repeated A legacy-encoded file read as UTF-8 (invalid byte sequences)
Cyrillic or CJK symbols in the middle of English words, like M黮ler A Western file read as a multi-byte Asian encoding
The same garbage getting longer each time Text that was mis-decoded and then saved again — double encoding

The crucial rule: don’t save a file while it looks garbled. Saving writes the wrong characters permanently. Close the file and reopen it with the right encoding.

Why old text files are still a problem

Text files, subtitles and e-books from the 2000s were usually saved in the system’s legacy code page, because that was Notepad’s default. Modern apps and phones assume UTF-8, so these files often open as gibberish.

The TXT to EPUB, TXT to DOCX and CSV to XLSX converters detect the encoding in this order:

  1. If the file starts with a BOM, follow it.
  2. If the bytes are valid UTF-8, use UTF-8.
  3. Otherwise try CP949 (Korean), Shift_JIS (Japanese), GB18030 (Chinese) and Windows-1251 (Cyrillic), and pick the one whose result reads as natural text in that script.
  4. Fall back to Windows-1252 (Western European).

If detection ever guesses wrong, you can pick the encoding manually under Options.

How to check a file’s encoding

  • Windows Notepad: the status bar in the bottom-right shows UTF-8, UTF-8 with BOM or ANSI (your system’s legacy code page).
  • VS Code: the encoding is shown in the bottom-right; click it to reopen the file with a different encoding.

Summary

Garbled text means the saving rule and the reading rule differ. Save new files as UTF-8, add a BOM to CSVs meant for Excel, and convert old legacy-encoded files once with an encoding-aware tool so they display correctly everywhere.

Tools used in this guide

Related guides