← XeFM crftwr/xefm on GitHub · craftware

Path Polymorphism System

Overview

The Path Polymorphism System is XeFM’s core abstraction layer that enables storage-agnostic code throughout the application. By extending the PathImpl interface with strategic virtual methods, the system eliminates all storage-specific conditionals from UI and dialog code, making it trivial to add new storage types without modifying existing code.

Architecture

Component Hierarchy

flowchart TB
    subgraph UI["UI / Dialog Layer — storage-agnostic (zero if/elif on storage type)"]
        direction LR
        TV["TextViewer"]
        FLM["FileListManager"]
        EP["ExternalPrograms"]
        FO["FileOperationService"]
    end

    Path["Path — facade (xefm.path)<br/>delegates every call to self._impl"]
    Impl["PathImpl — abstract base (xefm.path)<br/>Operations (abstract): exists · is_dir · iterdir · stat · open · rename · …<br/>Declarations: SCHEME · IS_REMOTE · CAPABILITIES · SEARCH_STRATEGY · DISPLAY_PREFIX<br/>Read through: get_scheme() · is_remote() · supports(name)<br/>Per instance: get_display_title() · get_extended_metadata()"]
    Local["LocalPathImpl<br/>xefm.path"]
    SSH["SSHPathImpl<br/>xefm.ssh"]
    S3["S3PathImpl<br/>xefm.s3"]
    Archive["ArchivePathImpl<br/>xefm.archive"]

    UI -->|polymorphic methods only| Path
    Path -->|delegates| Impl
    Impl --> Local & SSH & S3 & Archive

    classDef ui fill:#1e7e34,stroke:#7fd39b,color:#fff;
    classDef facade fill:#1a5490,stroke:#7fb3d5,color:#fff;
    classDef abc fill:#5e2d70,stroke:#b98fd0,color:#fff;
    classDef impl fill:#9a6308,stroke:#e0b45f,color:#fff;
    class TV,ID,SD,FO ui;
    class Path facade;
    class Impl abc;
    class Local,SSH,S3,Archive impl;

Design Principles

  1. Open/Closed Principle: Open for extension (new storage types), closed for modification (UI code)
  2. Dependency Inversion: UI depends on abstractions (PathImpl), not concrete implementations
  3. Single Responsibility: Each PathImpl subclass handles only its storage type
  4. Polymorphism Over Conditionals: Behavior varies through method overriding, not if/else checks

Path Facade and Migration History

The Facade / Implementation Split

The polymorphism system is built on a three-layer split in xefm/path.py:

_create_implementation is the real entry point for backend selection: it asks xefm.path_schemes which backend claims the URI’s scheme and falls back to LocalPathImpl when none does. Path.__init__ asks the same module whether the string is a URI at all, so that it survives instead of being run through PathlibPath — see “The Scheme Registry” below.

pathlib Compatibility

The facade preserves 100% compatibility with pathlib.Path:

Migration History

XeFM originally used pathlib.Path directly throughout src/. To make room for non-local storage without rewriting call sites, the codebase was migrated to the Path facade above. All src/ modules that touch paths were switched to from xefm.path import Path — including xefm/app.py, xefm/file_operations.py, xefm/pane_manager.py, xefm/state_manager.py, xefm/config.py, xefm/text_viewer.py, and the various dialog modules.

What early design notes framed as “future remote storage” now exists: the archive (xefm.archive.ArchivePathImpl), S3 (xefm.s3.S3PathImpl), and SSH/SFTP (xefm.ssh.SSHPathImpl) backends are all implemented and selected automatically by _create_implementation. Additional schemes (FTP, WebDAV, etc.) can be added the same way — see “Adding New Storage Types” below.

What a Backend Does, and What a Backend Is

PathImpl splits in two. The operations — exists, is_dir, iterdir, stat, open, rename, … — stay abstract: they are code, and only the backend can write them. What a backend is — its scheme, whether it is remote, what it can be asked to do, how its content is best searched — is not code. For every backend XeFM ships it is a compile-time constant, so it is declared rather than implemented.

The Five Declarations

class S3PathImpl(PathImpl):
    SCHEME = 's3'
    IS_REMOTE = True
    CAPABILITIES = frozenset({'write_operations', 'extraction_for_reading',
                              'cache_for_search'})
    SEARCH_STRATEGY = 'buffered'
    DISPLAY_PREFIX = 'S3: '
attribute type read through default
SCHEME str get_scheme() '' — every concrete backend must set it
IS_REMOTE bool is_remote() False
CAPABILITIES set of names supports(name) empty — nothing is assumed
SEARCH_STRATEGY 'streaming' / 'buffered' / 'extracted' the attribute 'buffered'
DISPLAY_PREFIX str, with trailing space the attribute ''

KNOWN_CAPABILITIES in xefm/path.py is the vocabulary:

name means
write_operations copy, move, create, delete
directory_rename renaming a directory, not just a file
file_editing handing the file to an external editor
streaming_read readable line by line, without a full fetch
extraction_for_reading content must be fetched or extracted first
cache_for_search that fetch is expensive enough to keep

As Shipped

  Local S3 SSH Archive
SCHEME file s3 ssh archive
IS_REMOTE False True True per path
write_operations ● ● ●  
directory_rename ●   ●  
file_editing ●      
streaming_read ●      
extraction_for_reading   ● ● ●
cache_for_search   ● ● ●
SEARCH_STRATEGY streaming buffered buffered extracted
DISPLAY_PREFIX '' 'S3: ' 'SSH: ' 'ARCHIVE: '

ArchivePathImpl.is_remote() is the one entry in this table that is still a method. An archive’s remoteness is its container’s: archive:// over a file on disk is local, the same URI over an object in S3 is not. It overrides the method and ignores IS_REMOTE; that is what the method is for.

Why Declarations and Not Methods

Before #426 each of these was an @abstractmethod that every backend answered with a one-line return. Twelve of them, four backends — and ten of the twelve had no caller anywhere in the application. One abstract declaration, four implementations, one façade delegation, a handful of tests, and nothing that read the answer. Publishing that interface would have meant asking an outside implementer ten questions and discarding every answer.

Three things follow from declaring them instead:

Reading Them Back

path.get_scheme()                    # 's3'
path.is_remote()                     # True
path.supports('write_operations')    # True
path.supports('directory_rename')    # False
path._impl.SEARCH_STRATEGY           # 'buffered'
path._impl.DISPLAY_PREFIX            # 'S3: '

Path delegates get_scheme, is_remote and supports; those have callers. SEARCH_STRATEGY and DISPLAY_PREFIX are read off the implementation class, and deliberately have no façade accessor: adding one is a two-line change for whoever finds the first real caller, and until then the codebase does not carry a method that nothing calls. The eight façade methods this replaced (supports_directory_rename, supports_file_editing, supports_write_operations, requires_extraction_for_reading, supports_streaming_read, should_cache_for_search, get_search_strategy, get_display_prefix) are gone; test/test_path_capabilities.py fails if one comes back.

The Two That Stayed Methods

get_display_title() and get_extended_metadata() compute something per instance, so they remain methods — but they are no longer abstract. Their defaults are str(self) and {'type': SCHEME, 'details': [], 'format_hint': 'standard'}, which is what every shipped backend except LocalPathImpl was already returning by hand. An outside implementation overrides them if it has something to show, and is not asked a question otherwise.

Two Base Classes, and Which One to Use

xefm/path_base.py holds two classes between PathImpl and a backend:

PathImpl (strict ABC, 48 abstract methods)
  └─ UriPathImpl          27 methods of string arithmetic; 14 still abstract
       └─ ReadOnlyPathImpl  the 11 writes refuse;  5 still abstract

UriPathImpl fills in everything that is string arithmetic. Every URI-addressed backend splits the same way — a root naming the container, a /-separated key inside it — so the class keeps one invariant:

str(self) == self._root_prefix() + self._key

and derives parent, parents, parts, name, stem, suffix, anchor, joinpath, with_name, with_stem, with_suffix, relative_to, glob, rglob, match, __eq__, __hash__, __lt__ and the rest from it. A subclass overrides at most two hooks — _parse(uri) (store what names the root, return the key) and _root_prefix() — and gets the arithmetic.

ReadOnlyPathImpl adds a policy, not more shared code: the eleven write operations raise OSError(errno.EROFS, …) and the capability declaration says so in advance. This is the class a virtual folder inherits.

Keeping the two apart matters because UriPathImpl is not read-only-specific. s3.py and ssh.py carry roughly 290 lines between them implementing the same 17 arithmetic methods twice — a debt UriPathImpl can absorb, and could not if the same class had also decided that writing raises.

The Cost of a Backend

base what it asks for
PathImpl 48 abstract methods
UriPathImpl 14 — the 5 below, plus the 9 write operations
ReadOnlyPathImpl 5 — exists, is_dir, iterdir, stat, open
from xefm.path_base import ReadOnlyPathImpl, UriStatResult

class RegistryPathImpl(ReadOnlyPathImpl):
    SCHEME = 'reg'

    def exists(self): ...
    def is_dir(self): ...
    def iterdir(self): ...           # yield self._child(name)
    def stat(self): ...              # return UriStatResult(size=…, mtime=…, is_dir=…)
    def open(self, mode='r', buffering=-1, encoding=None,
             errors=None, newline=None): ...

UriStatResult is there because all three shipped backends hand-roll a stat-result class — S3StatResult, a local class inside SSHPathImpl.stat, ArchiveEntry.to_stat_result. A virtual folder should not have to.

test/test_mock_storage_extensibility.py is the measure: its MockPathImpl was 384 lines and 61 methods on PathImpl, and is 85 lines and 15 on UriPathImpl. The read-only backend beside it is 25 lines and 5.

Why the Published Base Is Not PathImpl

PathImpl has never changed a signature. It has grown: +8 methods in v1.0.1, +1 in v1.0.4. The failure mode for a backend defined outside this repository is not a changed signature, it is a new @abstractmethod — and the moment one lands, every external subclass raises TypeError at import. For a class defined in ~/.xefm/config.py, that means the user’s entire configuration stops loading.

So the published extension point is ReadOnlyPathImpl, where a method that arrives carrying a default is invisible to code outside. The intermediate class is the compatibility buffer; the ergonomics are a side effect.

test/test_published_base.py enforces it: it asserts that ReadOnlyPathImpl.__abstractmethods__ is exactly those five names, so a new abstract method that reaches PathImpl without a default here fails the suite rather than a user’s config.

De-abstracting PathImpl itself would be the one-class version of this, and it would cost the check that catches an internal backend forgetting a method at import time. The ABC stays strict; leniency belongs in the published subclass.

Arithmetic Stays Inside the Class

UriPathImpl._at(key) builds a sibling through Path._from_impl(type(self)(…)) rather than Path(uri). Path(uri) asks the scheme factory which class to build, so a backend the factory has never heard of would find its own parent coming back as a local path. Going through the implementation it already has keeps the arithmetic closed over the class that produced it — and keeps this layer testable without the scheme registry existing at all.

The Scheme Registry

xefm/path_schemes.py is the one list of URI schemes XeFM understands. Before it, the same knowledge lived in four places that had to agree:

  what it decided
Path.__init__ whether a string survives verbatim instead of going through PathlibPath
Path._create_implementation which backend to build
XeFMApp._REMOTE_SCHEMES what Jump to Path resolves as a URI, and what a subshell refuses
config._REMOTE_SCHEMES what FAVORITE_DIRECTORIES / DRIVE_LOCATIONS list without probing

Patch only the first two — which is what a monkeypatch naturally does — and Jump to Path treats the new URI as a path relative to the current directory and reports “Path does not exist”. That report is accurate.

from xefm import path_schemes

path_schemes.register('reg', RegistryPathImpl, source='user')
path_schemes.is_uri('reg://HKEY_CURRENT_USER')   # True, everywhere
path_schemes.unregister_source('user')           # a config reload's other half
function  
register(scheme, factory, *, source) factory(uri) -> PathImpl
unregister_source(source) drop one source’s registrations, keep the rest
prefixes() ('archive://', 's3://', …), for a startswith test
is_uri(text) / scheme_of(text) what the four call sites above ask
create(path_str) the backend, or None so the caller falls back to local
validate_scheme(scheme) None, or why the name cannot be a scheme

The contract is the one the if-chain already kept: a URI string goes in, a PathImpl comes out, with the parsing inside the class. Nothing here looks at a URI beyond its scheme, which is what lets a backend be written and tested with this module nowhere in sight.

Registration is lazy. A factory imports its backend the first time it is asked for one, so listing the schemes — which the favourites picker does on every open, to decide whether a row needs probing — never pulls in boto3 for an s3:// row nobody selected. test_path_schemes.py checks that in a subprocess.

source groups registrations the way actions.registry does, so a config reload drops its own schemes without touching the built-ins. A registration that shadows another keeps what it covered: overriding s3:// and then reloading gives s3:// back rather than taking it away, and re-registering the same source over itself — which is what a reload is — does not stack.

scp:// and ftp://

Both sat in all three scheme lists with no branch in the factory, so they resolved to a local path: Path('ftp://h/p') became PathlibPath('ftp:/h/p') — the // collapsed — which then reported itself missing, naming a path the user had not typed.

They are registered to UnsupportedPathImpl (in xefm/path_base.py, and the smallest backend in the repository at five methods and no state). The URI stays intact, get_scheme() says 'ftp', is_remote() says True so a subshell still refuses it, exists() says False so a probe reports a missing path rather than raising — and anything that tries to read one gets OSError(ENOSYS, 'ftp:// is not supported by this version of XeFM').

Adding a Storage Type Inside This Repository

A backend that ships here subclasses PathImpl directly, declares the five attributes, implements the operations, and is registered in xefm/path_schemes.py for its scheme. Inheriting the strict ABC is the point: a missing method should fail at import for code in this repository.

class CustomPathImpl(PathImpl):
    """A backend for custom://resource/path."""

    SCHEME = 'custom'
    IS_REMOTE = True
    CAPABILITIES = frozenset({'extraction_for_reading', 'cache_for_search'})
    SEARCH_STRATEGY = 'buffered'
    DISPLAY_PREFIX = 'CUSTOM: '

    def __init__(self, uri: str):
        self._uri = uri
        # parse the URI here; the contract is "URI string in, PathImpl out"

    # … then the abstract operations: __str__, __eq__, __hash__, __lt__,
    # name/stem/suffix/parent/parts, exists, is_dir, iterdir, stat, open, …

PathImpl is a strict ABC on purpose: every operation stays abstract so that a backend inside this repository which forgets one fails at import, not in a pane. That same strictness is why PathImpl is not the class to hand to code outside this repository — see “Why the Published Base Is Not PathImpl” above.

The one thing the ABC cannot check is a declaration: a class attribute left at its default is not a missing method. test/test_path_capabilities.py carries that check instead — it fails when a shipped backend leaves SCHEME empty, and it holds the matrix above so a change to any row has to be a change to the table too.

Validation

PathImpl.__init_subclass__ checks the declarations as each subclass is defined, and warns and ignores rather than raising — the same treatment EVENT_HOOKS gives an unknown event name:

Warn-and-ignore is the direction of compatibility that matters here: a backend written against a later XeFM keeps the names this version understands and loses only the one it could not have acted on anyway.

Testing

file what it holds
test/test_path_capabilities.py the capability matrix, the validation rules, the list of retired methods
test/test_published_base.py the arithmetic, the read-only policy, and the five-method promise
test/test_path_schemes.py that one registration reaches all four call sites
test/test_mock_storage_extensibility.py a whole backend, written twice — writable and read-only — to prove the UI needs no change for either

References