← XeFM crftwr/xefm on GitHub · craftware

Text Encodings

The text viewer detects a file’s character encoding automatically, so legacy Japanese text — Shift-JIS, EUC-JP, ISO-2022-JP — displays correctly alongside UTF-8, with or without a BOM. When detection gets a file wrong, you can pick the encoding explicitly from inside the viewer.

Automatic detection

Opening a file in the text viewer (V) just works for:

The header’s right side names the encoding in use (e.g. Shift-JIS 12/340), so you always see what detection chose. Detection is content-based — no extension lists, no locale guessing — and something always displays: a file nothing else matches falls back to Latin-1 rather than refusing to open.

The same detection feeds the diff viewer, so comparing two Shift-JIS files shows real text on both sides. Content search (find_in_files) honors BOMs too: a UTF-16/32 file is grepped as text instead of being skipped as binary, and a UTF-8 BOM never blocks a ^-anchored match on the first line.

Choosing an encoding manually

Run change_encoding in the viewer to open the encoding picker:

↑/↓ choose — or just start typing (s jumps to Shift-JIS, eu to EUC-JP; the match is shown at the dialog’s bottom). Enter applies — the file is re-decoded in place, keeping your scroll position — and Esc cancels. The selection opens on whatever is currently in effect.

Plain E (edit_file) also works inside the viewer: it opens the viewed file in your configured editor, exactly as it does from the file list, and the viewer re-reads the file when a terminal editor returns.

A manual choice always shows something: bytes the chosen encoding can’t represent appear as replacement marks (�) instead of refusing. Choosing Auto returns to detection.

Two situations where the override earns its place:

The choice is per-file and per-viewing — closing the viewer forgets it, and reopening the file detects afresh.

Configuring the picker

The picker’s encoding list lives in ~/.xefm/config.py:

TEXT_ENCODINGS = ['utf-8', 'cp932', 'euc-jp', 'iso-2022-jp', 'latin-1']

Any Python codec name works — add 'koi8-r', 'gb2312', 'utf-16-le', or anything else you deal with. The list only feeds the manual picker; automatic detection is built in and unaffected.

The key is rebindable via KEY_BINDINGS['change_encoding'] like any other action (see Key Bindings).

See also