← XeFM crftwr/xefm on GitHub · craftware

Directory Scan System

Listing a pane means answering four questions about every entry: is it a directory, is it a symlink, how big is it, when was it modified — plus a fifth the details dialog needs, how much space does it take up. Asking the OS per file costs one round trip per file. On a local disk that is free. On a network mount it is the entire cost of the listing.

Measured on a 1,680-file SMB directory from a Synology NAS (~9 ms RTT), cold cache:

Approach Time
readdir alone 13 ms
getattrlistbulk — one bulk enumeration 2.0 s
os.scandir + one stat per entry 20.8 s
iterdir + 4 attribute calls per entry (the old path) 24.9 s

Finder lists the same directory in about 3 s, so ~2 s is close to the floor this NAS imposes; the old path sat roughly 8× above it.

The cost model, which the measurements match to within 1%:

1,680 × 12.4 ms ≈ 20.8 s, plus 5,040 × 0.8 ms ≈ 4 s of redundant calls, ≈ 24.8 s against 24.9 s measured.

The consequence is the important part: the redundancy was only ~17% of the cost. Deduplicating four calls per entry down to one still leaves 20.8 s. What matters is not touching files individually at all — SMB2 already returns size, mtime and type alongside the directory listing, and the old path discarded them and asked again per file.

Source: xefm/dir_scan.py, PathImpl.listdir_attrs. Tests: test/test_dir_scan.py, test/test_hidden_files.py.


The attribute record

Every backend returns the same record per entry, so callers never branch on platform:

{'is_dir': bool, 'is_link': bool, 'size': int, 'alloc': int | None,
 'mtime': float, 'hidden': bool, 'ok': bool}

is_dir, size and mtime describe the target of a symlink; is_link describes the link itself. That is exactly what stat()/is_dir() and is_symlink() reported when callers asked per file, so nothing downstream changed meaning. ok is False when the target could not be stat’d — a broken symlink — and the pane renders it as a link with --- for size and date. hidden is the exception to ok: it describes the directory entry itself, so it stays valid for an entry whose target is gone.

hidden is the platform’s mark, not the dot convention — see hidden entries below.

Directories report size: 0 on every backend. The bulk syscall cannot supply a directory’s size, and nothing displays or sorts on it (a directory renders as <DIR> and sorts as 0), so the backends are normalised rather than left to disagree. alloc follows it to 0 for the same reason — there is no directory equivalent of ATTR_FILE_ALLOCSIZE.

size and alloc

size is how many bytes long the entry is; alloc is how many bytes of the volume it occupies. They are Finder’s Size and on disk, and they part company for a sparse, compressed or cloned file: one Docker.raw under ~/Library/Containers is 994 GB long and 24 GB on disk. Summing the first and calling it disk usage is issue #275.

alloc is None when the backend cannot answer, which is not the same as 0:

Backend alloc Why
macOS ATTR_FILE_ALLOCSIZE, in the same bulk record free — one more attribute in a syscall already being made
Linux, Windows scandir st_blocks × 512 off the cached stat free — the stat is already taken
Windows None os.stat_result has no st_blocks; GetCompressedFileSize would be a call per file
SSH None ls -la reports a length; ls -s would be a second listing
S3, archive None no such concept; their hand-built stat_results carry st_blocks as None

st_blocks counts 512-byte units by definition, not the filesystem’s block size, so the multiplier is fixed. dir_scan.alloc_from_stat() is the single place that reads the field, and it maps both ways of it being missing — absent attribute, or present-but-None on a stat_result a backend built from a plain tuple — onto None.

Adding ATTR_FILE_ALLOCSIZE to the macOS request moved every later offset in the packed record: the kernel packs attributes in bitmap order, and ALLOCSIZE (bit 2) lands ahead of DATALENGTH (bit 9). A directory’s record stops before the file attributes altogether — it is shorter by exactly those two off_t — so they are read only for non-directories.

Per platform

Platform Mechanism Per-entry cost
macOS getattrlistbulk(2) — what Finder itself uses none
Windows os.scandir — DirEntry already carries the enumeration’s attributes none
Linux, other os.scandir — d_type answers is_dir/is_symlink, stat is cached one stat

Linux has no portable bulk equivalent, so it keeps one stat per entry — still a 4–6× reduction in calls, and local disks were never the problem. The win where it matters (macOS to a NAS, the case in issue #183) is the bulk syscall.

Symlinks are the one thing bulk enumeration cannot answer: its record describes the link, so dir_scan follows each one individually. A directory of symlinks costs what it always did; ordinary directories cost one scan.

getattrlistbulk failing on a volume that does not support it (some FUSE and network filesystems) falls back to os.scandir for that directory. Real directory errors — missing, permission denied, not a directory — still raise, so callers keep the error handling they had when this was iterdir.

Where it plugs in

PathImpl.listdir_attrs() returns [(Path, attrs), …] — iterdir plus everything a listing needs, in one call.

The default implementation is iterdir + per-entry stat, so backends with no bulk form (S3, SSH, archives) keep working untouched; LocalPathImpl overrides it with dir_scan.scan_dir. S3 already caches its list_objects_v2 response, which carries Size and LastModified, so the bulk-metadata idea has precedent — the local backend was the one asking per file.


Hidden entries

A leading dot hides an entry on POSIX. Windows says the same thing with a file attribute and no dot, and XeFM only ever tested the name — so with hidden files off, a Windows pane still listed AppData, $Recycle.Bin, System Volume Information and every desktop.ini (issue #284).

The attribute is part of the record because reading it is free where it exists: DirEntry.stat(follow_symlinks=False) on Windows is served from the enumeration, so the one-pass scan already has it. Elsewhere hidden is False and costs nothing to fill in.

Two predicates in dir_scan put the two conventions together, and every consumer of the hidden-files toggle calls one of them:

Predicate For a caller holding Cost
is_hidden(name, attrs) a scan record none
is_hidden_path(path) only a path one lstat, Windows only

FILE_ATTRIBUTE_HIDDEN alone decides it. FILE_ATTRIBUTE_SYSTEM on its own sits on folders a user still expects to see — C:\Windows\Fonts, a customized Documents — while everything Explorer calls a protected operating system file carries HIDDEN as well as SYSTEM. Testing HIDDEN therefore catches pagefile.sys and System Volume Information without swallowing Fonts.

hidden is read from the entry, never from a symlink’s target: a followed stat describes something else, so attrs_via_path and attrs_for_path take a second lstat for links only.

macOS has UF_HIDDEN, which Finder honours; dir_scan does not read it. The bulk record would need ATTR_CMN_FLAGS, which shifts every offset in _parse_bulk_entry, and nothing on macOS commonly carries the flag.


Sorting and filtering without re-reading

The second half of issue #183, and the one users actually felt. Sorting a directory needs no information the listing did not already collect, yet every sort and filter change used to re-read the whole directory — paying the full cost above to produce a list the pane could already derive.

FileListManager now splits the listing in two:

Method Reads the disk? What it does
compute_listing(path, …) yes, once scan the directory, then assemble
_assemble_listing(entries, …) no filter, sort, build the display cache
recompute_listing(pane, …) no re-assemble from the pane’s snapshot

apply_listing stores the scan in pane['_listing_entries'], and XeFMApp._resort rebuilds from it on the current tick. A re-sort of the 1,680- file NAS directory costs microseconds instead of a full re-read.

The snapshot is taken after the hidden-file filter and before the filename filter. So:

_resort falls back to _relist whenever the pane has no snapshot — nothing listed yet, or the last listing failed — so behaviour is unchanged when there is nothing to reuse. A failed listing clears the snapshot rather than leaving a later sort to re-filter a directory that can no longer be read.

Two user-visible consequences

The cursor now follows the file, not the row. Re-sorting used to leave focused_index where it was, so the cursor landed on whatever file happened to occupy that row. _resort keeps it on the same file. Filter changes still reset the cursor to the top, which is what set_filter has always done.

No blank-then-repopulate. _relist clears pane['files'] until a worker reports back, which on a slow mount left the pane empty for the whole re-read. An in-memory re-sort lands on the same tick, so the list never empties.

What a sort no longer picks up

A sort used to re-read the directory, so it incidentally refreshed the pane. It no longer does. External changes arrive through FileMonitorManager, which watches the directory and posts a reload — that is the mechanism responsible for freshness, and it is unaffected. Every post-operation path still re-lists for real. The trade-off is only visible where file monitoring is unavailable and the directory changed underneath: previously a sort would have surfaced it by accident, now it waits for an explicit refresh.


Other readers of the snapshot

Sorting was the first consumer of pane['_listing_entries'], not the only one. Any feature that wants is_dir / size / mtime for entries a pane has already listed should read them from there rather than ask the filesystem again — otherwise it reintroduces exactly the per-file round trips this system exists to remove, on top of a listing that already paid for the answers.

Compare & Select (issue #245)

compute_compare_selection joins the two panes by name and tests each pair against the enabled relations. It used to call is_dir() on every entry of both sides for the join, then stat() on both halves of every matched pair — two per-file calls per entry, per side, for every criteria, including the pure filename join that reads no attributes at all. Two 1,680-entry panes meant ~6,700 calls, or roughly 45 s under the cost model above, for a comparison whose every input the panes already held.

It now takes the records instead:

compute_compare_selection(current_files, other_files, criteria,
                          current_attrs=..., other_attrs=...)   # {str(path): record}

XeFMApp._pane_attrs(pane) builds one side’s mapping from _listing_entries; both the inline path and the content-comparison worker pass it. Measured on two 200-file directories, all relations: 1,200 per-file calls → 0.

Three rules this follows, and any future consumer should:

The staleness trade-off is the one _resort already made: a comparison now reflects the same snapshot the pane’s own size and date columns are drawn from, so it agrees with what the user is looking at, and is refreshed by the same mechanisms. Content comparison is unaffected in kind — no snapshot can answer “are these bytes equal”, so it still reads both files on the task worker; only the size short-circuit that decides whether to read now comes for free.