Canonical developer reference for XeFM’s archive support. Two independent paths:
xefm/archive.py
(ArchivePathImpl + handlers + cache), plugged into the Path abstraction.xefm/app.py (the XeFMApp create/extract methods), using
the stdlib zipfile / tarfile modules directly, and falling back to the
registered handler for anything they cannot read (§1.1).Which formats are readable is decided at import, not written down. Reading
goes through a registry (§1.1) whose libarchive-backed entries depend on what the
library that actually loaded can do, so anything enumerating formats has to be
generated from archive_readable_formats() rather than kept as a list
somewhere. The Help dialog’s “Archive Formats” table is the one place that
enumerates them for a user; XeFMApp._archive_help_section() builds it, pairing
that call with _writable_formats() for the “Create” column.
Source of truth is the code; this document summarizes structure and intent, not every line.
Browsing an archive works because xefm/archive.py implements the PathImpl
interface, so archive contents flow through the same Path machinery as local
and S3 paths.
archive://<absolute_path_to_archive>#<internal_path>
archive:///home/user/data.zip# (archive root)
archive:///home/user/data.zip#folder/ (a directory inside)
archive:///home/user/data.zip#folder/file.txt (a file inside)
The # separates the archive file path from the internal path. Path() detects
the archive:// scheme and constructs an ArchivePathImpl (xefm/path.py).
A @dataclass giving a uniform view of an entry across formats: name,
internal_path, is_dir, size, compressed_size, mtime, mode,
archive_type. Helpers:
to_stat_result() — an os.stat_result so archive entries interoperate with
filesystem-shaped code.from_zip_info(zip_info, archive_type='zip') / from_tar_info(tar_info,
archive_type='tar') — classmethod factories from zipfile.ZipInfo /
tarfile.TarInfo.ArchiveHandler is the base interface for reading an archive: open(),
close(), list_entries(internal_path=""), get_entry_info(internal_path),
extract_to_bytes(internal_path), extract_to_file(internal_path, target_path),
iter_member_bytes(internal_path, chunk_size), entry_count(),
iter_extract(dest_dir, password=None), encryption_status(),
verify_password(pwd), plus context-manager support. The last five are what a
third format needed and the first two did not have: extraction and encryption
used to be answered by asking whether the handler was a ZipHandler, and
reading one member was a single opaque call (§1.3).
_build_index(entries, archive_type) on the base class fills _entry_cache and
_directory_cache and synthesizes a virtual directory entry for every parent an
archive names only implicitly. iter_extract has a generic implementation there
too, walking that index one entry at a time and refusing members whose path
escapes the destination (is_safe_member_path).
Three concrete handlers exist:
ZipHandler — ZIP via zipfile. Caches entries on open, with lazy loading
for large archives (>1000 entries: only shallow structure is cached up front,
deeper entries load on demand via getinfo). Keeps its own copy of the
indexing loop rather than calling _build_index, because that lazy policy is
zip-only. Also carries the encryption read path (see §3).TarHandler(archive_path, compression=None) — tar and compressed variants
(gz, bz2, xz) via tarfile, indexed through _build_index.LibarchiveHandler(archive_path, label) — everything libarchive
contributes: .7z, .rar, .iso, .cab, .cpio, .rpm (§1.2).All three download a remote archive (is_remote()) to a temp file on open()
and delete it on close().
ARCHIVE_HANDLERS is a list of ArchiveFormat(label, suffixes, factory,
description), and it is the single answer to “can XeFM read this file”. It
replaced an if/elif chain in ArchiveCache._create_handler plus two
isinstance(handler, ZipHandler) tests in the password gate.
| Function | Answers |
|---|---|
register_archive_format(fmt) |
add, replacing any entry with the same label |
archive_format_for_name(name) |
the matching ArchiveFormat, or None |
archive_format_label(name) |
its label — 'zip', 'tar.gz', '7z' |
archive_strip_suffix(name) |
the name with its archive suffix removed |
archive_readable_formats() |
every registered format, for enumeration |
archive_writable_formats() |
the formats that brought a writer with them |
Three rules hold for anything registered:
.tar.gz before .tar right only because of where the branches sat in the
source; matching now sorts by suffix length, and a test pins it both ways round.ArchiveFormat.writer, set only where
the writer arrives with the engine that reads it. zip and tar leave it None
and are created by XeFMApp through zipfile / tarfile, so “what can P create”
is the union of two sources (§2) — libarchive reads strictly more formats than
it writes (rar, lha and cab are read-only), so the two lists cannot be one.Registration happens at import: _register_builtin_formats() at the bottom of
xefm/archive.py for zip and tar, then register_libarchive_formats() for
whatever the loaded library justifies. xefm/archive_libarchive.py imports
xefm/archive.py in turn; the cycle resolves because the registration call sits
below every name it needs.
xefm/archive_libarchive.py)libarchive-c is a pure-ctypes binding that carries no binary, so the shared
library comes from one of three places, in this order: the LIBARCHIVE
environment variable, a bundled copy, or find_library("archive") —
the system copy. Nothing is required: with no usable library the registry simply
has fewer entries and zip and tar are unaffected.
The bundled copy is found by bundled_library_path(), which looks in
xefm/_bin/ — inside XeFM’s own package, which is the one directory this module
can locate from __file__ without knowing anything about the bundle around it.
A source checkout leaves it empty; windows_app/build.ps1 fills it in the
copied package (§ “Step 4b” of
WINDOWS_APP_BUILD_SYSTEM.md) with a DLL downloaded
from crftwr/xefm-bin-deps, pinned by
release tag and SHA-256.
Finding one is how LIBARCHIVE comes to be set when the user did not set it:
libarchive-c reads that variable once, at import, and offers no other way to
choose a library, so _use_bundled_library() puts the path there before
_probe() imports the binding. That is also why this lives in the module that
does the import rather than in XeFM’s startup, where it would be one
import-order mistake away from having no effect. A LIBARCHIVE the user set
themselves is left alone — naming a library is answering exactly this question.
Because two of those three paths are built by someone else, capability is
probed, never inferred from a version. archive_version_details() names the
codecs actually compiled in; _CANDIDATES says what each format needs, and a
format registers only when the library exports every reader symbol it lists,
every filter symbol, and reports every codec.
| Label | Suffixes | Needs | Writer |
|---|---|---|---|
7z |
.7z |
7zip reader, liblzma | 7zip, compression=lzma2 |
rar |
.rar |
rar and rar5 readers | — |
iso |
.iso |
iso9660 reader | iso9660 |
cab |
.cab |
cab reader, zlib (MSZIP is deflate) | — |
cpio |
.cpio |
cpio reader | cpio_newc, hdrcharset=UTF-8 |
rpm |
.rpm |
cpio reader, rpm filter, zlib + liblzma | — |
RAR requires both generations because a .rar is RAR4 or RAR5 and offering the
suffix on one reader would be a lie for half of them; the win over rarfile is
that libarchive implements them itself, with none of the non-free unrar binary.
cpio_newc rather than the plain cpio writer: the historic odc format stores
sizes in eight octal digits and so cannot hold a member over 8 GB.
Probing is also what keeps the silent external-program fallback out of reach:
libarchive answers a missing stream codec by spawning gzip -d or zstd -d, one
process per archive, and on Windows those binaries do not exist. This is not
hypothetical. macOS’s system libarchive reports no libzstd and still reads a
.tar.zst — with PATH emptied it admits why:
ArchiveError Can't initialize filter; unable to run program "zstd -d -qq"
It had been shelling out to Homebrew’s zstd. A format whose codec is missing
must never be offered, which is why .tar.zst goes through the standard library
instead (§2) and not through here.
The fallback is a filter mechanism, not a format one, and the difference
matters because it is easy to over-generalize the paragraph above. A codec named
by an entry inside a format is decoded by that format’s reader, which cannot
reach __archive_read_program at all. Measured on the same library, with the
zstd binary present on PATH:
| Archive | Result |
|---|---|
| a 7z whose entry is zstd-compressed | ZSTD codec is unsupported — a clean failure, no process |
a .tar.zst |
unable to run program "zstd -d -qq" — the fallback |
So a codec the library lacks costs a spawned process only where it applies to a
stream: a bare compressed file, or a payload the filter chain sees. .rpm is
the one registered format where that is reachable, since its payload is a
filtered stream — nothing outside the file says which filter, so the probe cannot
refuse it in advance. Everything an entry inside a 7z, RAR, ISO, CAB or cpio
declares fails visibly instead, which is the behaviour we want and the reason the
probe only has to cover what a format needs to open at all.
The three builds are not the same library, and one difference is visible.
macOS 3.7.4 system zlib liblzma bz2lib
Linux 3.8.x distro zlib liblzma bz2lib liblz4 libzstd (+ openssl, expat, …)
Windows 3.8.9 bundled zlib liblzma bz2lib libzstd cng libb2
Every one of them registers the same six formats, because the candidates need
only zlib and liblzma. What differs is what an entry inside one can be
compressed with: a zstd-compressed 7z member reads on Windows and Linux and
fails on macOS with ZSTD codec is unsupported. That is the honest answer — the
per-entry codec is named inside the file and cannot be probed at registration —
but it does mean the same archive can open on one machine and not another, which
is worth recognising in a bug report rather than rediscovering. The Windows
build’s cng is not a fourth difference: it spares the DLL an OpenSSL
dependency and decrypts nothing, since through 3.8.9 zip is the only format
libarchive decrypts at all.
register_libarchive_formats() says something only when the outcome is not the
ordinary one. Success is debug — a full line at every startup was more than the
normal case deserved, and the Help dialog’s “Archive Formats” table carries the
same information where a user can find it. What stays visible is the two
outcomes worth acting on: no library at all, and a library that loaded and then
justified nothing, which is the shape a mis-built or half-stripped copy takes
and is easy to mistake for the first.
No random access. libarchive is a forward stream of headers: open() makes
one pass to build the index, and every later read re-opens the file and scans to
its entry. Browsing suits that (the structure is cached once), but extracting n
entries one at a time is O(n²) on a solid archive — which is why
LibarchiveHandler overrides iter_extract with a single pass, and why
entry_count() is overridden too: that pass yields the archive’s stored
members, not the directories the index invented for them.
Writing is write_archive(archive_path, sources, format_name=…, options=…,
on_entry=…, on_bytes=…), registered as the 7z format’s writer when the library
exports archive_write_set_format_7zip. member_walk() produces members
depth-first, a directory before its children, deliberately matching
_count_archive_entries(include_dirs=True) member for member — including
counting an unlistable directory as itself and not descending — because that
pass’s total is the one the write has to reach. Directories are stored rather
than implied, so an empty one survives. options='compression=lzma2' is
explicit: libarchive’s 7z writer defaults to LZMA1, while 7-Zip itself has
written LZMA2 for years. libarchive’s 7z writer has no encryption, so XeFM
cannot create a password-protected 7z.
Progress, both directions. Neither path uses libarchive’s own
archive_read_extract_set_progress_callback: XeFM does not use
archive_read_extract at all, writing the blocks itself, which is what makes
block-level granularity available for free. On extraction iter_extract yields
each entry before writing its payload and calls on_bytes(n) per block; on
creation write_archive calls on_entry(arcname, size) before each member and
on_bytes(n) as the source is read. Both feed the same
:class:~xefm.archive_progress.ByteProgress the stdlib paths use — see the note
at the end of §2.
Encryption is two questions, not one. Which entries are encrypted comes from
archive_entry_is_encrypted on the headers read at open(). Whether they can be
decrypted at all is can_decrypt_7z(), which decrypts a 183-byte AES-256 7z
embedded in the module.
The answer today is no, everywhere: through 3.8.9 libarchive decrypts ZIP and
no other format, and its 7z reader rejects an encrypted entry outright — not
for want of a crypto library, which was the first and wrong reading of macOS’s
crypto-less 3.7.4. A 3.8.9 Windows build linked against CNG refuses the same
archive, which is what settled it. The probe is kept for two reasons anyway: it
is the difference between “not supported” and rejecting every password the user
types, and the day a release does add 7z decryption it starts returning True with
no change here. archive_version_details() could not answer this either way —
it names codecs, not what the format readers do with them.
ArchiveCache(max_open=5, ttl=300) keeps recently used handlers open so
repeated navigation doesn’t re-open the archive each time:
max_open handlers are live.ttl seconds is closed on next
access.threading.RLock.get_stats() (open_archives, cache_hits, cache_misses,
hit_rate, evictions, avg_open_time, …)._create_handler is a lookup in the registry (§1.1) — archive_format_for_name
then fmt.factory(archive_path) — raising ArchiveFormatError when nothing
registered reads the name. A process-wide instance is returned
by get_archive_cache(), which reads ARCHIVE_CACHE_MAX_OPEN /
ARCHIVE_CACHE_TTL from config (falling back to 5 / 300).
ArchivePathImpl(archive_uri, metadata=None) implements PathImpl for archive
members: URI parsing, path properties (name, stem, suffix, parent,
parts, …), path manipulation (joinpath, with_name, relative_to, …),
queries (exists, is_dir, is_file, stat), directory traversal (iterdir,
glob, rglob), and read-only I/O (open, read_text, read_bytes). All
write/mutate operations (write_*, mkdir, unlink, rename, chmod, …)
raise OSError("Archive files are read-only").
extract_to_stream(stream, progress_callback) is the copy-out path (§1.3).
It declares what an archive is rather than answering it method by method:
SCHEME = 'archive', CAPABILITIES = {'extraction_for_reading',
'cache_for_search'} (no write capability — archives are read-only here),
SEARCH_STRATEGY = 'extracted', and get_extended_metadata() for the info
dialog. is_remote() stays a method, because an archive’s remoteness is its
container’s. See doc/dev/PATH_POLYMORPHISM_SYSTEM.md. A per-instance _property_cache
memoizes name / parts; a _metadata['entry'] slot caches the resolved
ArchiveEntry.
libarchive keeps an entry’s pathname in both a wide and a narrow form, and
converts between them using a code page it gets by calling
setlocale(LC_CTYPE, NULL) in its own C runtime — get_current_codepage() in
archive_string.c. On macOS and Linux that resolves to UTF-8 and none of this is
visible. On Windows it is the ANSI code page, 1252 on a US install, and any name
1252 cannot spell makes archive_entry_pathname() return NULL. What that does
depends on who asked:
Pathname required and fails the whole archive;Two things answer this, and they are separate because they fix different halves.
_use_utf8_ctype() puts the process’s C locale on UTF-8 (Windows only, before
the binding is imported). Python’s own encodings are untouched — it derives those
from GetACP(), not from the C locale — so the only code this reaches is
libarchive’s conversions. This is also why crftwr/xefm-bin-deps links the
shared MSVC runtime: a statically linked one is private to archive.dll and
cannot be reached from here at all.
That fixes ISO but not cpio, whose default is the OEM code page (437), which
libarchive derives from a table of locale names and which the .UTF8 suffix
therefore does not move. So cpio is told its charset outright, on both sides —
write_options='hdrcharset=UTF-8' and the matching _CHARSET_BY_LABEL entry
that _open_reader() reads.
_open_reader() exists only for that: hdrcharset has to be set between
archive_read_new and archive_read_open, and libarchive-c’s file_reader
does both in one call. It rebuilds the reader from the same pieces, and falls
back to plain file_reader if that package is ever rearranged.
_CHARSET_BY_LABEL deliberately holds only cpio and rpm. Forcing UTF-8 on a
CAB whose names are CP932 does not garble them — it makes every entry’s pathname
NULL, so the archive opens and looks empty. Mojibake is a bad listing; nothing
at all is a broken one, and libarchive’s own default is the better answer for
every format that stores a legacy code page.
Path.copy_to has a branch for archive → file that streams the member into
the destination through ArchivePathImpl.extract_to_stream, which walks
iter_member_bytes() and calls the progress callback per block. Without it the
copy fell into copy_to’s generic arm — read_bytes() then write_bytes() —
and that one opaque call cost three things at once: the whole member in memory,
no byte bar, and no cancellation, because for a cross-storage copy
FileOperationService._remote_progress puts task.checkpoint() inside the
progress callback and nothing ever called it. A large file inside a 7z is where
that is unmissable, but zip and tar behaved identically.
LibarchiveHandler.iter_member_bytes coalesces libarchive’s own ~16 KiB blocks
up to chunk_size (1 MiB, matching file_operations._CHUNK): the consumer takes
a lock on the UI’s progress state per block, and a gigabyte at 16 KiB would do
that sixty thousand times.
A cancel raised inside the callback propagates unchanged — copy_to guards the
callback so it comes back as the caller’s own exception rather than an OSError
about a failed copy — and the branch removes the truncated destination on the way
out.
Still generic:
archive→s3/ssh. Those combinations have no branch and fall through toread_bytes()/write_bytes(), so copying a member straight from a browsed archive to remote storage still buffers it and still cannot be cancelled. The fix is the same shape as the local one, needing a file-like adapter overiter_member_bytes()forupload_from_stream.
xefm/app.py handles entering an archive: when the cursor is on a recognized archive
file and Enter is pressed, it remembers the cursor and sets the pane path to
Path(f"archive://{entry.absolute()}#"). Because ArchivePathImpl.parent of the
archive root is the archive file’s containing directory, “up” exits the archive
naturally. Nested archives (an archive inside a browsed archive) are not
supported.
A small exception hierarchy under ArchiveError (each carries a technical
message and a user-facing user_message): ArchiveFormatError,
ArchiveCorruptedError, ArchiveExtractionError, ArchiveNavigationError,
ArchivePermissionError, ArchiveDiskSpaceError, plus the encryption pair
ArchivePasswordRequired and ArchiveEncryptionUnsupported (§3).
ArchiveCache is lock-guarded; handlers are read-only and independent. Multiple
threads may read the same or different archives concurrently through the cache.
Archives are never modified while open.
Creation and extraction are not in xefm/archive.py — they live on XeFMApp
in xefm/app.py and operate on local filesystem paths using the stdlib directly.
There is no separate ArchiveOperations/ArchiveUI class.
Creation has two implementations, so “what can P create” has two sources. The
stdlib half is class data on XeFMApp:
_ARCHIVE_EXTS — the extensions zipfile / tarfile can create → format label._TAR_MODES — format label → tarfile write mode (w, w:gz, w:bz2,
w:xz, w:zst); ZIP is handled separately.Both grow a Zstandard row when xefm.archive.tar_zstd_supported() is true —
'zst' in TarFile.OPEN_METH, which is Python 3.14 and up. The readable registry
applies the same condition, so .tar.zst is creatable exactly when it is
openable. Zstandard deliberately does not come from libarchive: see the
external-program evidence in §1.2.
The other half is the registry’s writers (§1.1). _writable_formats() is their
union, sorted longest-suffix-first — sorted rather than concatenated for the
same reason the read registry sorts, so .tar.gz beats .tar whichever list
each came from. On top of it:
_archive_format(name) → the label P can create, or None._readable_archive_format(name) → the label Enter browses and U extracts._archive_basename(name) → archive_strip_suffix(name), the default
extraction subdirectory.A name the registry reads but brought no writer for is refused by P with a
message saying so, rather than silently gaining a .tar.gz suffix.
_add_to_zip(zf, path, arcname, task=None, prog=None, bytes_=None) — adds a
path to an open ZipFile, recursing into directories (zipfile, unlike tarfile,
does not recurse on its own)._write_archive(sources, archive_path, fmt, task=None, prog=None) — writes
sources into a new archive. ZIP uses ProgressZipFile(..., "w", ZIP_DEFLATED);
tar formats use ProgressTarFile.open(..., _TAR_MODES[fmt]) (tarfile recurses
into directories, and its add(filter=…) hook is where per-member progress and
cancellation are taken). Both subclasses are stdlib passthroughs that count
bytes — see The byte bar below. Returns the number of files added._entry_size(path) — a source file’s size for the byte bar, 0 when unreadable
(the writer is what reports a genuinely broken file)._count_archive_entries(sources, include_dirs) — the counting pass that makes
the progress bar determinate: leaf files only for ZIP, files and directories
for tar, matching what each writer actually adds. An unreadable directory counts
as itself and is not descended.create_archive() (the P key) — the UI flow: takes the active pane’s
selection (or the focused entry), prompts for a filename, and writes the
archive into the other pane’s directory. An unrecognized extension defaults
to .tar.gz. A single selected item prefills "<basename>.". Overwrite is
confirmed via a message box. Guards refuse to archive entries that live inside a
read-only archive or to write into one._extract_archive(archive_path, dest_dir, fmt, pwd=None, task=None, prog=None)
— extracts into dest_dir (created if absent) and returns the entry count. Tar
extraction uses the filter="data" argument where available (Python 3.12+) to
reject unsafe member paths, falling back when the argument is unsupported. pwd
is the password for an encrypted ZIP, verified up front (see §3) so a wrong
password fails before any file is written._reporting_members(members, describe, task, prog, bytes_=None) — the generator
handed to extractall(members=…); describe(member) yields its (name, size).
Progress and cancellation are taken per member yielded rather than by
hand-rolling the loop, so extractall’s deferred directory-permission fix-up (a
read-only directory would otherwise block writing into it) and zipfile’s member
path sanitization both still run.extract_archive() (the U key) — the UI flow: takes the active pane’s
selection (or the focused entry) like every other file operation
(_selected_or_focused; U ignoring the selection was issue #408) and extracts
each archive into a subdirectory named after it (_archive_basename) in the
other pane’s directory. Selected entries that are not readable archives are
dropped from the plan and reported as a count — a single target keeps its own
diagnosis (“is not a file” / “is not a supported archive”), which a count would
throw away. Confirms when CONFIRM_EXTRACT_ARCHIVE is set or a destination
already exists. Refuses nested archives and extracting into a read-only
archive.Both flows hand the work to xefm.task rather than doing it in the dialog
callback that started it (issue #280) — compressing a large tree, or building a
compressed tar’s member list, would otherwise block the event loop for its whole
duration: no repaint, no keys, no way out.
_submit_archive_task(task, run, on_done, dest_dir) is the shared submit. Each
flow builds a Task (kind="archive_create" / "archive_extract", progress
started as OperationType.ARCHIVE_CREATE / ARCHIVE_EXTRACT) whose run body
counts, then writes or extracts, and returns its outcome as a dict — {"added":
n} for a create, {"done": archives, "entries": n, "failures": [(entry, exc)]}
for an extract (plus "cancelled": True and the half-written "partial" when it
was stopped), or {"error": exc} — which on_done reports on the main thread.
A whole batch of archives extracts inside one run: start_operation is
re-issued per archive (the bar restarts, named after the archive), task.title
carries “archive i of N”, and one archive’s failure is collected rather than
raised, so the rest of the batch still runs. The submit also brackets dest_dir’s
filesystem watcher for the run, the same suppression copy/move/delete use so an
operation’s own writes don’t re-list the watching pane throughout (issue #243).
This buys the standard ProgressDialog: a determinate items bar, the current
entry’s name, and Esc to cancel. Cancellation unwinds from the per-entry
checkpoint; a cancelled create deletes the half-written archive (it either did
not exist before, or an overwrite truncated it the moment the file opened), while
a cancelled extract leaves what landed — the destination may be a directory the
user already had files in, so removing it wholesale could take those with it.
Extraction’s failure dispatch is ordered most-specific-first, because
NotImplementedError is a RuntimeError: encryption XeFM cannot decrypt is
reported as unsupported, and only a plain RuntimeError counts as a wrong
password. That dispatch lives on the worker now (see below), so a wrong password
re-asks in place instead of unwinding the task and resubmitting it.
xefm/archive_progress.py)The dialog’s secondary bar shows the current member’s bytes, the same meaning it has for a copy — without it a single large member (a VM image, a video) leaves the item bar still for minutes.
The payload copy is buried inside zipfile / tarfile, and rewriting those loops
to count bytes would mean re-deriving each member’s metadata and, on the extract
side, their safety checks: zipfile’s path sanitization and tarfile’s sparse-file
and deferred directory-permission handling. So ProgressZipFile /
ProgressTarFile instead override the one method the payload actually flows
through and wrap the file object passing by. The stdlib loop runs untouched:
| Operation | Seam | Counts |
|---|---|---|
| zip create | ZipFile.open(zinfo, 'w') — write() copies into it |
writes |
| zip extract | ZipFile.open(member) — _extract_member copies out of it |
reads |
| tar create | TarFile.addfile(tarinfo, fileobj) — add() hands it the source |
reads |
| tar extract | TarFile.makefile — proxy swapped over self.fileobj for the call |
reads |
ByteProgress holds the current member’s total (from stat(), ZipInfo.file_size
or TarInfo.size) and rate-limits reports by volume — at most ~200 per member,
never oftener than every 64 KiB — so an 8 KiB-chunked gigabyte does not cost a
hundred thousand lock acquisitions. start() must be called after
update_progress, which clears the byte fields for the incoming item. With no
ByteProgress attached both classes are pure passthroughs.
Measured overhead of the whole task path, 1500 files: ~1.06x for zip, within noise
for .tar.xz (the counting pass is 11 ms of it).
Create: ZIP, TAR, TAR.GZ (.tgz), TAR.BZ2 (.tbz2), TAR.XZ (.txz) and — on
3.14+ — TAR.ZST (.tzst) from the _ARCHIVE_EXTS table, all of it stdlib, plus
.7z, .iso and .cpio from the registry’s writers. .rar, .cab and .rpm
are readable and not creatable. Ask _writable_formats().
Extract and browse: those, plus whatever libarchive contributed (§1.2) — .7z
where a usable library loaded. Ask archive_readable_formats(); do not restate
the list.
Both _extract_archive and _write_archive route a format that is neither
"zip" nor in _TAR_MODES away from the stdlib: to _extract_via_handler,
which drives the handler’s iter_extract, and to _write_via_handler, which
calls the registry entry’s writer. Both supply the per-entry bookkeeping the
stdlib paths get from _reporting_members / the tar filter=report hook —
task.checkpoint(), prog.update_progress(name), bytes_.start(size) — and
hand bytes_.advance down as the block callback.
Single-file gzip/bzip2/xz streams are readable as members but are not first-class create targets in the flow above.
The create/extract flow works on local filesystem paths and does not perform cross-storage staging. (Remote-archive support exists only on the read/browse side, where a handler downloads the archive to a temp file.)
Password-protected ZIP support (extract and browse). Python’s zipfile decrypts
only legacy ZipCrypto; WinZip AES (compression method 99) cannot be
decrypted and is detected and refused with a clear message. No third-party
dependency (pyzipper etc.) is used.
xefm/archive.py)A module-level dict keyed by the archive file’s absolute path, guarded by a lock,
holding passwords for the session (in-memory only, nothing persisted):
set_archive_password, get_archive_password, clear_archive_password.
The handler contract speaks a format-neutral vocabulary —
encryption_status() → 'none' | 'password' | 'unsupported' — because 7z is
routinely encrypted and the gate could not go on naming zip’s schemes. The
zip-level names survive one level down, inside ZipHandler:
zip_encryption_status(zf) → 'none' | 'zipcrypto' | 'aes' (AES wins if any
entry uses it). ZipHandler.encryption_status() maps zipcrypto → 'password'
and aes → 'unsupported'.archive_encryption_status_path(path) — the same classification from a file
path, for a raw file rather than a browsed handler. It goes through the
registry (so any format can answer) but not through ArchiveCache (extraction
is not browsing).archive_extraction_survey(path) → (status, members, bytes) — that same open
asked for the progress totals as well, and what the extract flow actually
calls. members / bytes come from ArchiveHandler.extraction_totals(),
which counts what an extraction walks: the archive’s stored members, not the
browsable index, whose invented parent directories no member matches. A file
that will not open comes back ('none', 0, 0) — missing from the bar rather
than promised to it, so the bar cannot end up stuck short of full.verify_zip_password(zf, pwd) — opens the smallest encrypted entry to validate
the ZipCrypto header cheaply. No-op when nothing is encrypted; raises
RuntimeError (missing/wrong password) or NotImplementedError (AES).extract_to_bytes / extract_to_file pass pwd=get_archive_password(...) to
ZipFile.read, mapping RuntimeError → ArchivePasswordRequired (via
_read_runtime_error) and NotImplementedError → ArchiveEncryptionUnsupported.
encryption_status() and verify_password(pwd) expose the helpers per handler.
Thin wrappers so the app never reaches into _impl / cache internals:
get_member_archive_path(path) — the archive file behind an archive://
member Path, else None.archive_password_state(path) → 'ok' | 'need' | 'unsupported'. Ordinary paths return
'ok' cheaply (nothing opened), so every read can route through it.try_archive_password(path, password) — verify (UTF-8 encoded) and, on success,
remember it; returns a bool.xefm/app.py)Extract — extract_archive surveys the whole batch first, on the worker
(_survey_archives → archive_extraction_survey per archive, each a
cancellation point): one open per archive answers its encryption status and its
progress totals together. The probe is a full open of the file, which is why it
is not done on the main thread; the totals ride along on an open that had to
happen anyway.
The batch then runs as one progress operation, started with the task and
never restarted — it used to be started afresh per archive, so the bar rewound
to zero at every archive boundary — with update_operation_total(items,
total_bytes=...) published once for the lot. Each archive’s own extraction is
passed owns_total=False so it does not rescale that bar to its own size.
The workers are one pool for the whole batch (_ExtractionPool), started
once and parked between archives rather than joined at each boundary. That
boundary used to cost the tail of every archive: its last members’ closes ran
with the other workers already idle, and the next archive could not begin
reading until they finished — on a destination that uploads at close, the
expensive part. The batch’s own thread still publishes one archive at a time
and keeps the password gate, the title and per-archive attribution; what it
waits for is that archive’s reading, not its writing. Concurrency stays at
the operation’s worker count throughout, since a boundary that briefly doubled it
would be a worse bargain than the one it removed.
Members are closed one at a time, through the same
file_operations.CloseGate a copy uses — the constraint belongs to the
destination, not to the operation writing it. See
PARALLEL_COPY_IMPLEMENTATION.md for the reasoning and the numbers.
Counts and failures therefore arrive late: an archive’s last files land after
its reading ended, so neither its member count nor its failure exists when the
loop moves on. close() drains the pool and returns {key: (written, error)},
which run folds into done / entries / failures after the loop.
Each archive is read by _extract_members: symmetric workers
(transfer_workers on the destination’s scheme) each take one member and carry it from
claim to closed file, so a member’s whole life — and its file handle’s — stays
on one thread. The archive is claimed one member at a time, but that lock
lives with the cursor it protects, inside
LibarchiveHandler.extraction_pass: _extract_members never learns why the
claims serialise. What overlaps is everything after the claim, which on a
filesystem that holds a file in a local cache until it is closed (WebDAV,
NFS’s close-to-open flush) is the whole transfer — two workers halved a
batch’s wall clock against one such mount. On a filesystem that does not, the
workers queue on the claim and the run is what it always was; the design does
not have to know which it is dealing with. The worker count cannot be derived
from the destination — a mounted volume reports the same file scheme a local
disk does — which is why it is a setting.
Progress there goes through the copy engine’s transfer slots, not the single
current-item fields: several members are in flight, so there is no one current
member to name. file_end counts the item after the file is closed, so a
member counts when it has actually landed, and file_closing marks the row
for the duration of that close — on a destination that uploads there, it is
most of the member’s time and was previously shown as nothing at all. The zip
and tar paths cannot do the same: extractall opens and closes each
destination file itself, and reimplementing its loop to get at that seam was
ruled out (see archive_progress).
Both flows also log at the grain a copy does — one line per file, none for
a directory: Extracted 'sub/b.txt': photos.7z → /dest/photos and
Added 'src/a.txt' → backup.7z. Where a seam exists after the member (the zip
create loop; the extraction workers, which own theirs to the closed file) the
line is written there. zipfile.extractall, tarfile.add and libarchive’s
writer drive their own member loop and offer a seam only at the start of one,
so those lines are held one step behind by _StepBehindLog: a member that then
fails is never claimed as written. Extraction workers share one sink, so the
flow wraps it in file_operations.serialized_log — the same wrapper, for the
same reason, as the parallel copy.
On the surveyed status: 'unsupported' is recorded as a failure for that
archive; 'password' asks for one through the task’s UI bridge (Task.ask, the
same seam the copy conflict dialog uses — the masked prompt stacks at z + 5,
above the progress dialog); anything else extracts directly. The up-front
verify_zip_password means a wrong password re-asks with an error and never
leaves a half-extracted directory; cancelling the prompt cancels the batch. A
working password is stored so a later browse reuses it.
Browse / view — _ensure_archive_password gates opening a file that may
live in an encrypted ZIP: 'ok' runs the open callback immediately; 'aes'
shows a message; 'need' shows a masked prompt, verifies via
try_archive_password, and re-prompts on failure. Listing an encrypted archive
needs no password (the ZIP central directory is unencrypted); the prompt is
deferred to the first file open.
The password prompt is a masked field. xefm/input_dialog.py’s
show_input(..., password=True) forwards mask="•" to PuiKit’s TextEdit, whose
masking is length-preserving (cursor/selection/hit-test still map onto the real
buffer) and disables copy/cut so plaintext never reaches the clipboard. The
widget itself lives in the PuiKit repo (puikit/widgets/text_edit.py).
# xefm/_config.py
ARCHIVE_CACHE_MAX_OPEN = 5 # max archives kept open by the browse cache
ARCHIVE_CACHE_TTL = 300 # cache TTL in seconds
CONFIRM_EXTRACT_ARCHIVE = True # confirm before extracting
# Key bindings
'create_archive': {'keys': ['P'], 'selection': 'required'}
'extract_archive': ['U']
test/test_archive_*.py — entry conversion, handlers, cache (LRU/TTL), and
ArchivePathImpl.test/test_archive_path_impl.py — ArchivePathImpl, plus (in
TestStreamingOutOfAnArchive) the copy-out path of §1.3: that
iter_member_bytes really streams for zip and tar, that copy_to reports more
than once, and that a cancel raised from the callback propagates and leaves no
partial file.test/test_archive_registry.py — the readable-format table: longest-suffix
matching both ways round, replacement by label, dispatch through
_create_handler, the base class’s unencrypted defaults, is_safe_member_path,
and the generic iter_extract.test/test_archive_libarchive.py — the 7z path, skipped wholesale where no
usable libarchive exists: browsing, extraction, creation, the count agreeing
with the counting pass, both progress bars moving (a multi-block member has to
report more than once and land on full), and cancellation mid-archive.
Fixtures are written with libarchive’s own writers, so no external 7z binary
is needed; the encrypted one is a stored blob, because libarchive cannot
write an encrypted 7z. ISO and cpio round-trip through XeFM’s own create and
extract, with a long name and a non-ASCII one in the tree because plain ISO
9660 would mangle both and only Rock Ridge / Joliet keep them. RAR, CAB and RPM
have no content fixture — libarchive writes none of them and there is no
way to generate one on the test machine — so what is pinned is their
registration and their read-only property; that a real .rar opens is a
hand-check. Its encryption assertions
branch on can_decrypt_7z(), which is False on macOS’s system library — so the
correct-password case is exercised only where a crypto-capable build is loaded.test/test_archive_password.py — classification, verification, the registry,
the ZipHandler read path, and the gate helpers (hermetic base64 ZipCrypto
fixture).test/test_xefm_app_archive_password.py — _extract_archive and the extract UI
flow (prompt, wrong-then-right retry, AES refusal, plain zip), and
_ensure_archive_password.test/test_archive_task.py — the task path: per-entry progress and
cancellation in both loops, the counting pass agreeing with each writer, the
byte bar streaming inside a large member in all four directions (plus the
metadata and path-sanitization the stdlib hooks exist to preserve), what a
cancelled or failed create leaves on disk, and a real app on a MemoryBackend
running both flows on the worker.xefm/archive.py; Path factory in xefm/path.py.XeFMApp in xefm/app.py; byte counting in
xefm/archive_progress.py.xefm/task.py; the copy/move/delete equivalent in
xefm/file_operations.py.xefm/s3.py.