← XeFM crftwr/xefm on GitHub · craftware

Archive System

Canonical developer reference for XeFM’s archive support. Two independent paths:

Which formats are readable is decided at import, not written down. Reading goes through a registry (§1.1) whose libarchive-backed entries depend on what the library that actually loaded can do, so anything enumerating formats has to be generated from archive_readable_formats() rather than kept as a list somewhere. The Help dialog’s “Archive Formats” table is the one place that enumerates them for a user; XeFMApp._archive_help_section() builds it, pairing that call with _writable_formats() for the “Create” column.

Source of truth is the code; this document summarizes structure and intent, not every line.


1. Read / browse path (virtual directory)

Browsing an archive works because xefm/archive.py implements the PathImpl interface, so archive contents flow through the same Path machinery as local and S3 paths.

Archive URI format

archive://<absolute_path_to_archive>#<internal_path>

archive:///home/user/data.zip#                  (archive root)
archive:///home/user/data.zip#folder/           (a directory inside)
archive:///home/user/data.zip#folder/file.txt   (a file inside)

The # separates the archive file path from the internal path. Path() detects the archive:// scheme and constructs an ArchivePathImpl (xefm/path.py).

ArchiveEntry

A @dataclass giving a uniform view of an entry across formats: name, internal_path, is_dir, size, compressed_size, mtime, mode, archive_type. Helpers:

ArchiveHandler and subclasses

ArchiveHandler is the base interface for reading an archive: open(), close(), list_entries(internal_path=""), get_entry_info(internal_path), extract_to_bytes(internal_path), extract_to_file(internal_path, target_path), iter_member_bytes(internal_path, chunk_size), entry_count(), iter_extract(dest_dir, password=None), encryption_status(), verify_password(pwd), plus context-manager support. The last five are what a third format needed and the first two did not have: extraction and encryption used to be answered by asking whether the handler was a ZipHandler, and reading one member was a single opaque call (§1.3).

_build_index(entries, archive_type) on the base class fills _entry_cache and _directory_cache and synthesizes a virtual directory entry for every parent an archive names only implicitly. iter_extract has a generic implementation there too, walking that index one entry at a time and refusing members whose path escapes the destination (is_safe_member_path).

Three concrete handlers exist:

All three download a remote archive (is_remote()) to a temp file on open() and delete it on close().

1.1 The readable-format registry

ARCHIVE_HANDLERS is a list of ArchiveFormat(label, suffixes, factory, description), and it is the single answer to “can XeFM read this file”. It replaced an if/elif chain in ArchiveCache._create_handler plus two isinstance(handler, ZipHandler) tests in the password gate.

Function Answers
register_archive_format(fmt) add, replacing any entry with the same label
archive_format_for_name(name) the matching ArchiveFormat, or None
archive_format_label(name) its label — 'zip', 'tar.gz', '7z'
archive_strip_suffix(name) the name with its archive suffix removed
archive_readable_formats() every registered format, for enumeration
archive_writable_formats() the formats that brought a writer with them

Three rules hold for anything registered:

Registration happens at import: _register_builtin_formats() at the bottom of xefm/archive.py for zip and tar, then register_libarchive_formats() for whatever the loaded library justifies. xefm/archive_libarchive.py imports xefm/archive.py in turn; the cycle resolves because the registration call sits below every name it needs.

1.2 The libarchive engine (xefm/archive_libarchive.py)

libarchive-c is a pure-ctypes binding that carries no binary, so the shared library comes from one of three places, in this order: the LIBARCHIVE environment variable, a bundled copy, or find_library("archive") — the system copy. Nothing is required: with no usable library the registry simply has fewer entries and zip and tar are unaffected.

The bundled copy is found by bundled_library_path(), which looks in xefm/_bin/ — inside XeFM’s own package, which is the one directory this module can locate from __file__ without knowing anything about the bundle around it. A source checkout leaves it empty; windows_app/build.ps1 fills it in the copied package (§ “Step 4b” of WINDOWS_APP_BUILD_SYSTEM.md) with a DLL downloaded from crftwr/xefm-bin-deps, pinned by release tag and SHA-256.

Finding one is how LIBARCHIVE comes to be set when the user did not set it: libarchive-c reads that variable once, at import, and offers no other way to choose a library, so _use_bundled_library() puts the path there before _probe() imports the binding. That is also why this lives in the module that does the import rather than in XeFM’s startup, where it would be one import-order mistake away from having no effect. A LIBARCHIVE the user set themselves is left alone — naming a library is answering exactly this question.

Because two of those three paths are built by someone else, capability is probed, never inferred from a version. archive_version_details() names the codecs actually compiled in; _CANDIDATES says what each format needs, and a format registers only when the library exports every reader symbol it lists, every filter symbol, and reports every codec.

Label Suffixes Needs Writer
7z .7z 7zip reader, liblzma 7zip, compression=lzma2
rar .rar rar and rar5 readers —
iso .iso iso9660 reader iso9660
cab .cab cab reader, zlib (MSZIP is deflate) —
cpio .cpio cpio reader cpio_newc, hdrcharset=UTF-8
rpm .rpm cpio reader, rpm filter, zlib + liblzma —

RAR requires both generations because a .rar is RAR4 or RAR5 and offering the suffix on one reader would be a lie for half of them; the win over rarfile is that libarchive implements them itself, with none of the non-free unrar binary. cpio_newc rather than the plain cpio writer: the historic odc format stores sizes in eight octal digits and so cannot hold a member over 8 GB.

Probing is also what keeps the silent external-program fallback out of reach: libarchive answers a missing stream codec by spawning gzip -d or zstd -d, one process per archive, and on Windows those binaries do not exist. This is not hypothetical. macOS’s system libarchive reports no libzstd and still reads a .tar.zst — with PATH emptied it admits why:

ArchiveError Can't initialize filter; unable to run program "zstd -d -qq"

It had been shelling out to Homebrew’s zstd. A format whose codec is missing must never be offered, which is why .tar.zst goes through the standard library instead (§2) and not through here.

The fallback is a filter mechanism, not a format one, and the difference matters because it is easy to over-generalize the paragraph above. A codec named by an entry inside a format is decoded by that format’s reader, which cannot reach __archive_read_program at all. Measured on the same library, with the zstd binary present on PATH:

Archive Result
a 7z whose entry is zstd-compressed ZSTD codec is unsupported — a clean failure, no process
a .tar.zst unable to run program "zstd -d -qq" — the fallback

So a codec the library lacks costs a spawned process only where it applies to a stream: a bare compressed file, or a payload the filter chain sees. .rpm is the one registered format where that is reachable, since its payload is a filtered stream — nothing outside the file says which filter, so the probe cannot refuse it in advance. Everything an entry inside a 7z, RAR, ISO, CAB or cpio declares fails visibly instead, which is the behaviour we want and the reason the probe only has to cover what a format needs to open at all.

The three builds are not the same library, and one difference is visible.

macOS   3.7.4  system      zlib liblzma bz2lib
Linux   3.8.x  distro      zlib liblzma bz2lib liblz4 libzstd (+ openssl, expat, …)
Windows 3.8.9  bundled     zlib liblzma bz2lib libzstd cng libb2

Every one of them registers the same six formats, because the candidates need only zlib and liblzma. What differs is what an entry inside one can be compressed with: a zstd-compressed 7z member reads on Windows and Linux and fails on macOS with ZSTD codec is unsupported. That is the honest answer — the per-entry codec is named inside the file and cannot be probed at registration — but it does mean the same archive can open on one machine and not another, which is worth recognising in a bug report rather than rediscovering. The Windows build’s cng is not a fourth difference: it spares the DLL an OpenSSL dependency and decrypts nothing, since through 3.8.9 zip is the only format libarchive decrypts at all.

register_libarchive_formats() says something only when the outcome is not the ordinary one. Success is debug — a full line at every startup was more than the normal case deserved, and the Help dialog’s “Archive Formats” table carries the same information where a user can find it. What stays visible is the two outcomes worth acting on: no library at all, and a library that loaded and then justified nothing, which is the shape a mis-built or half-stripped copy takes and is easy to mistake for the first.

No random access. libarchive is a forward stream of headers: open() makes one pass to build the index, and every later read re-opens the file and scans to its entry. Browsing suits that (the structure is cached once), but extracting n entries one at a time is O(n²) on a solid archive — which is why LibarchiveHandler overrides iter_extract with a single pass, and why entry_count() is overridden too: that pass yields the archive’s stored members, not the directories the index invented for them.

Writing is write_archive(archive_path, sources, format_name=…, options=…, on_entry=…, on_bytes=…), registered as the 7z format’s writer when the library exports archive_write_set_format_7zip. member_walk() produces members depth-first, a directory before its children, deliberately matching _count_archive_entries(include_dirs=True) member for member — including counting an unlistable directory as itself and not descending — because that pass’s total is the one the write has to reach. Directories are stored rather than implied, so an empty one survives. options='compression=lzma2' is explicit: libarchive’s 7z writer defaults to LZMA1, while 7-Zip itself has written LZMA2 for years. libarchive’s 7z writer has no encryption, so XeFM cannot create a password-protected 7z.

Progress, both directions. Neither path uses libarchive’s own archive_read_extract_set_progress_callback: XeFM does not use archive_read_extract at all, writing the blocks itself, which is what makes block-level granularity available for free. On extraction iter_extract yields each entry before writing its payload and calls on_bytes(n) per block; on creation write_archive calls on_entry(arcname, size) before each member and on_bytes(n) as the source is read. Both feed the same :class:~xefm.archive_progress.ByteProgress the stdlib paths use — see the note at the end of §2.

Encryption is two questions, not one. Which entries are encrypted comes from archive_entry_is_encrypted on the headers read at open(). Whether they can be decrypted at all is can_decrypt_7z(), which decrypts a 183-byte AES-256 7z embedded in the module.

The answer today is no, everywhere: through 3.8.9 libarchive decrypts ZIP and no other format, and its 7z reader rejects an encrypted entry outright — not for want of a crypto library, which was the first and wrong reading of macOS’s crypto-less 3.7.4. A 3.8.9 Windows build linked against CNG refuses the same archive, which is what settled it. The probe is kept for two reasons anyway: it is the difference between “not supported” and rejecting every password the user types, and the day a release does add 7z decryption it starts returning True with no change here. archive_version_details() could not answer this either way — it names codecs, not what the format readers do with them.

ArchiveCache

ArchiveCache(max_open=5, ttl=300) keeps recently used handlers open so repeated navigation doesn’t re-open the archive each time:

_create_handler is a lookup in the registry (§1.1) — archive_format_for_name then fmt.factory(archive_path) — raising ArchiveFormatError when nothing registered reads the name. A process-wide instance is returned by get_archive_cache(), which reads ARCHIVE_CACHE_MAX_OPEN / ARCHIVE_CACHE_TTL from config (falling back to 5 / 300).

ArchivePathImpl

ArchivePathImpl(archive_uri, metadata=None) implements PathImpl for archive members: URI parsing, path properties (name, stem, suffix, parent, parts, …), path manipulation (joinpath, with_name, relative_to, …), queries (exists, is_dir, is_file, stat), directory traversal (iterdir, glob, rglob), and read-only I/O (open, read_text, read_bytes). All write/mutate operations (write_*, mkdir, unlink, rename, chmod, …) raise OSError("Archive files are read-only").

extract_to_stream(stream, progress_callback) is the copy-out path (§1.3).

It declares what an archive is rather than answering it method by method: SCHEME = 'archive', CAPABILITIES = {'extraction_for_reading', 'cache_for_search'} (no write capability — archives are read-only here), SEARCH_STRATEGY = 'extracted', and get_extended_metadata() for the info dialog. is_remote() stays a method, because an archive’s remoteness is its container’s. See doc/dev/PATH_POLYMORPHISM_SYSTEM.md. A per-instance _property_cache memoizes name / parts; a _metadata['entry'] slot caches the resolved ArchiveEntry.

Filenames on Windows

libarchive keeps an entry’s pathname in both a wide and a narrow form, and converts between them using a code page it gets by calling setlocale(LC_CTYPE, NULL) in its own C runtime — get_current_codepage() in archive_string.c. On macOS and Linux that resolves to UTF-8 and none of this is visible. On Windows it is the ANSI code page, 1252 on a US install, and any name 1252 cannot spell makes archive_entry_pathname() return NULL. What that does depends on who asked:

Two things answer this, and they are separate because they fix different halves.

_use_utf8_ctype() puts the process’s C locale on UTF-8 (Windows only, before the binding is imported). Python’s own encodings are untouched — it derives those from GetACP(), not from the C locale — so the only code this reaches is libarchive’s conversions. This is also why crftwr/xefm-bin-deps links the shared MSVC runtime: a statically linked one is private to archive.dll and cannot be reached from here at all.

That fixes ISO but not cpio, whose default is the OEM code page (437), which libarchive derives from a table of locale names and which the .UTF8 suffix therefore does not move. So cpio is told its charset outright, on both sides — write_options='hdrcharset=UTF-8' and the matching _CHARSET_BY_LABEL entry that _open_reader() reads.

_open_reader() exists only for that: hdrcharset has to be set between archive_read_new and archive_read_open, and libarchive-c’s file_reader does both in one call. It rebuilds the reader from the same pieces, and falls back to plain file_reader if that package is ever rearranged.

_CHARSET_BY_LABEL deliberately holds only cpio and rpm. Forcing UTF-8 on a CAB whose names are CP932 does not garble them — it makes every entry’s pathname NULL, so the archive opens and looks empty. Mojibake is a bad listing; nothing at all is a broken one, and libarchive’s own default is the better answer for every format that stores a legacy code page.

1.3 Copying a member out

Path.copy_to has a branch for archive → file that streams the member into the destination through ArchivePathImpl.extract_to_stream, which walks iter_member_bytes() and calls the progress callback per block. Without it the copy fell into copy_to’s generic arm — read_bytes() then write_bytes() — and that one opaque call cost three things at once: the whole member in memory, no byte bar, and no cancellation, because for a cross-storage copy FileOperationService._remote_progress puts task.checkpoint() inside the progress callback and nothing ever called it. A large file inside a 7z is where that is unmissable, but zip and tar behaved identically.

LibarchiveHandler.iter_member_bytes coalesces libarchive’s own ~16 KiB blocks up to chunk_size (1 MiB, matching file_operations._CHUNK): the consumer takes a lock on the UI’s progress state per block, and a gigabyte at 16 KiB would do that sixty thousand times.

A cancel raised inside the callback propagates unchanged — copy_to guards the callback so it comes back as the caller’s own exception rather than an OSError about a failed copy — and the branch removes the truncated destination on the way out.

Still generic: archive → s3 / ssh. Those combinations have no branch and fall through to read_bytes() / write_bytes(), so copying a member straight from a browsed archive to remote storage still buffers it and still cannot be cancelled. The fix is the same shape as the local one, needing a file-like adapter over iter_member_bytes() for upload_from_stream.

xefm/app.py handles entering an archive: when the cursor is on a recognized archive file and Enter is pressed, it remembers the cursor and sets the pane path to Path(f"archive://{entry.absolute()}#"). Because ArchivePathImpl.parent of the archive root is the archive file’s containing directory, “up” exits the archive naturally. Nested archives (an archive inside a browsed archive) are not supported.

Error handling

A small exception hierarchy under ArchiveError (each carries a technical message and a user-facing user_message): ArchiveFormatError, ArchiveCorruptedError, ArchiveExtractionError, ArchiveNavigationError, ArchivePermissionError, ArchiveDiskSpaceError, plus the encryption pair ArchivePasswordRequired and ArchiveEncryptionUnsupported (§3).

Thread safety

ArchiveCache is lock-guarded; handlers are read-only and independent. Multiple threads may read the same or different archives concurrently through the cache. Archives are never modified while open.


2. Create / extract path

Creation and extraction are not in xefm/archive.py — they live on XeFMApp in xefm/app.py and operate on local filesystem paths using the stdlib directly. There is no separate ArchiveOperations/ArchiveUI class.

Format detection

Creation has two implementations, so “what can P create” has two sources. The stdlib half is class data on XeFMApp:

Both grow a Zstandard row when xefm.archive.tar_zstd_supported() is true — 'zst' in TarFile.OPEN_METH, which is Python 3.14 and up. The readable registry applies the same condition, so .tar.zst is creatable exactly when it is openable. Zstandard deliberately does not come from libarchive: see the external-program evidence in §1.2.

The other half is the registry’s writers (§1.1). _writable_formats() is their union, sorted longest-suffix-first — sorted rather than concatenated for the same reason the read registry sorts, so .tar.gz beats .tar whichever list each came from. On top of it:

A name the registry reads but brought no writer for is refused by P with a message saying so, rather than silently gaining a .tar.gz suffix.

Creation

Extraction

Running as a task

Both flows hand the work to xefm.task rather than doing it in the dialog callback that started it (issue #280) — compressing a large tree, or building a compressed tar’s member list, would otherwise block the event loop for its whole duration: no repaint, no keys, no way out.

_submit_archive_task(task, run, on_done, dest_dir) is the shared submit. Each flow builds a Task (kind="archive_create" / "archive_extract", progress started as OperationType.ARCHIVE_CREATE / ARCHIVE_EXTRACT) whose run body counts, then writes or extracts, and returns its outcome as a dict — {"added": n} for a create, {"done": archives, "entries": n, "failures": [(entry, exc)]} for an extract (plus "cancelled": True and the half-written "partial" when it was stopped), or {"error": exc} — which on_done reports on the main thread. A whole batch of archives extracts inside one run: start_operation is re-issued per archive (the bar restarts, named after the archive), task.title carries “archive i of N”, and one archive’s failure is collected rather than raised, so the rest of the batch still runs. The submit also brackets dest_dir’s filesystem watcher for the run, the same suppression copy/move/delete use so an operation’s own writes don’t re-list the watching pane throughout (issue #243).

This buys the standard ProgressDialog: a determinate items bar, the current entry’s name, and Esc to cancel. Cancellation unwinds from the per-entry checkpoint; a cancelled create deletes the half-written archive (it either did not exist before, or an overwrite truncated it the moment the file opened), while a cancelled extract leaves what landed — the destination may be a directory the user already had files in, so removing it wholesale could take those with it.

Extraction’s failure dispatch is ordered most-specific-first, because NotImplementedError is a RuntimeError: encryption XeFM cannot decrypt is reported as unsupported, and only a plain RuntimeError counts as a wrong password. That dispatch lives on the worker now (see below), so a wrong password re-asks in place instead of unwinding the task and resubmitting it.

The byte bar (xefm/archive_progress.py)

The dialog’s secondary bar shows the current member’s bytes, the same meaning it has for a copy — without it a single large member (a VM image, a video) leaves the item bar still for minutes.

The payload copy is buried inside zipfile / tarfile, and rewriting those loops to count bytes would mean re-deriving each member’s metadata and, on the extract side, their safety checks: zipfile’s path sanitization and tarfile’s sparse-file and deferred directory-permission handling. So ProgressZipFile / ProgressTarFile instead override the one method the payload actually flows through and wrap the file object passing by. The stdlib loop runs untouched:

Operation Seam Counts
zip create ZipFile.open(zinfo, 'w') — write() copies into it writes
zip extract ZipFile.open(member) — _extract_member copies out of it reads
tar create TarFile.addfile(tarinfo, fileobj) — add() hands it the source reads
tar extract TarFile.makefile — proxy swapped over self.fileobj for the call reads

ByteProgress holds the current member’s total (from stat(), ZipInfo.file_size or TarInfo.size) and rate-limits reports by volume — at most ~200 per member, never oftener than every 64 KiB — so an 8 KiB-chunked gigabyte does not cost a hundred thousand lock acquisitions. start() must be called after update_progress, which clears the byte fields for the incoming item. With no ByteProgress attached both classes are pure passthroughs.

Measured overhead of the whole task path, 1500 files: ~1.06x for zip, within noise for .tar.xz (the counting pass is 11 ms of it).

Supported formats

Create: ZIP, TAR, TAR.GZ (.tgz), TAR.BZ2 (.tbz2), TAR.XZ (.txz) and — on 3.14+ — TAR.ZST (.tzst) from the _ARCHIVE_EXTS table, all of it stdlib, plus .7z, .iso and .cpio from the registry’s writers. .rar, .cab and .rpm are readable and not creatable. Ask _writable_formats().

Extract and browse: those, plus whatever libarchive contributed (§1.2) — .7z where a usable library loaded. Ask archive_readable_formats(); do not restate the list.

Both _extract_archive and _write_archive route a format that is neither "zip" nor in _TAR_MODES away from the stdlib: to _extract_via_handler, which drives the handler’s iter_extract, and to _write_via_handler, which calls the registry entry’s writer. Both supply the per-entry bookkeeping the stdlib paths get from _reporting_members / the tar filter=report hook — task.checkpoint(), prog.update_progress(name), bytes_.start(size) — and hand bytes_.advance down as the block callback.

Single-file gzip/bzip2/xz streams are readable as members but are not first-class create targets in the flow above.

The create/extract flow works on local filesystem paths and does not perform cross-storage staging. (Remote-archive support exists only on the read/browse side, where a handler downloads the archive to a temp file.)


3. Encryption

Password-protected ZIP support (extract and browse). Python’s zipfile decrypts only legacy ZipCrypto; WinZip AES (compression method 99) cannot be decrypted and is detected and refused with a clear message. No third-party dependency (pyzipper etc.) is used.

Password registry (xefm/archive.py)

A module-level dict keyed by the archive file’s absolute path, guarded by a lock, holding passwords for the session (in-memory only, nothing persisted): set_archive_password, get_archive_password, clear_archive_password.

Classification / verification helpers

The handler contract speaks a format-neutral vocabulary — encryption_status() → 'none' | 'password' | 'unsupported' — because 7z is routinely encrypted and the gate could not go on naming zip’s schemes. The zip-level names survive one level down, inside ZipHandler:

ZipHandler read path

extract_to_bytes / extract_to_file pass pwd=get_archive_password(...) to ZipFile.read, mapping RuntimeError → ArchivePasswordRequired (via _read_runtime_error) and NotImplementedError → ArchiveEncryptionUnsupported. encryption_status() and verify_password(pwd) expose the helpers per handler.

UI-facing gate helpers

Thin wrappers so the app never reaches into _impl / cache internals:

Flows (xefm/app.py)

Masked input (PuiKit)

The password prompt is a masked field. xefm/input_dialog.py’s show_input(..., password=True) forwards mask="•" to PuiKit’s TextEdit, whose masking is length-preserving (cursor/selection/hit-test still map onto the real buffer) and disables copy/cut so plaintext never reaches the clipboard. The widget itself lives in the PuiKit repo (puikit/widgets/text_edit.py).


Configuration

# xefm/_config.py
ARCHIVE_CACHE_MAX_OPEN = 5      # max archives kept open by the browse cache
ARCHIVE_CACHE_TTL      = 300    # cache TTL in seconds
CONFIRM_EXTRACT_ARCHIVE = True  # confirm before extracting

# Key bindings
'create_archive': {'keys': ['P'], 'selection': 'required'}
'extract_archive': ['U']

Tests

References