Platform & Infrastructure

2026-09-21

The Stateful Node: Architectural Vision for Sovereign Edge Infrastructure

A reference architecture for running stateful workloads on sovereign edge infrastructure without sacrificing cattle-not-pets principles. Design patterns for the four-tier storage hierarchy, workload-aware caching, and automated recovery.


Platform & Infrastructure · architecture vision — implementation in progress · ~9 min read · companion: The Second Life of Computers · cited by the Sovereign Fleet series

Where the Sovereign Edge Stops

The sovereign edge architecture handles stateless workloads well. Workers, KV, D1, R2 — those primitives can build a surprising amount from scratch. What they cannot run is an existing application: the database that needs POSIX filesystem semantics, the search index that needs low-latency block I/O, the dependency graph that nobody is going to rewrite as a serverless function.

Those workloads have to live somewhere. The question this post answers is: where can they run that is sovereign, reliable and cost-effective — without rebuilding the vendor lock-in the rest of the architecture exists to escape?

The Principle: A Recoverable Pet, Not True Cattle

The stateful node is the compatibility layer that completes the sovereign edge. It rests on three principles:

  1. Location-agnostic. Whether the machine sits in a garage, a data centre or a cloud region is irrelevant. The architecture cares only about reliability — can it survive failure and recover automatically?
  2. Sovereign by design. No vendor lock-in. Provider-agnostic abstractions, so moving from one provider to another — or on-premises — does not change the architecture.
  3. Cattle, with one honest exception. Machine death should be a five-minute inconvenience, not a disaster. For stateful data the honest description is narrower: this is a recoverable pet — 5–30 minutes of recovery, not true cattle. Saying so is the point; the alternative is claiming a property the architecture does not have.

The Architecture: A Four-Tier Storage Hierarchy

Each tier solves a different problem, and the boundary between them is the whole design.

TierWhat it isWhat it is for
1 · EdgeGlobally distributed, HTTP-accessible, no filesystem semanticsThe stateless application
2 · Network-attached blockDetachable block storage (a cloud provider's volumes)The data. Survives machine death; POSIX semantics; ~0.5–1 ms
3 · Local SSD cachebcache or lvmcache on local NVMe, transparent to the applicationPerformance. Local SSD is roughly an order of magnitude faster than the network volume
4 · Object backupEncrypted, deduplicated backups to object storageRecovery from volume loss; point-in-time restore

The critical design decision is the order. Cache sits between the application and the network volume — application → cache → (miss) → volume — never application → volume → cache. Get that backwards and the cache stops being a cache.

The Five Load-Bearing Decisions

Each is a trade-off made deliberately, and each has a failure mode attached.

1 · Durability: detachable volumes, disposable system disk

Problem: local disk dies with the machine, and then the data is a pet. Decision: network-attached volumes hold the data; the system disk holds only the OS and is reproducible from a script.

  • Machine death (≈5 minutes): provision a new machine, detach the volume from the dead one, attach it to the new one, start containers.
  • Volume loss (≈30 minutes): provision a new machine, create a fresh volume, restore from the object backup, start containers.

The system disk is disposable; the data volume is precious. Quarterly recovery drills are what keep that claim true.

2 · Performance: workload-aware caching, with the safe mode as the default

Local SSD cache in front of the network volume, configured per workload. The applications and kernel both cache, and they are not interchangeable:

LayerToolPurpose
ApplicationDatabase buffer poolPages the database manages itself
ApplicationRedisShared in-memory cache
KernelbcacheBlock-level cache in front of the network volume

bcache's mode is a data-loss decision, not a tuning knob:

ModeBehaviourRisk
writethroughWrites go to cache and backing volume togetherDefault. No data loss if the cache device dies; slower writes
writebackWrites land in cache first, flushed laterDangerous. Cache failure means data loss — only with a UPS, monitoring, and an explicit decision to accept it
writearoundWrites bypass the cache entirelySafe; cache serves reads only. Good for write-heavy workloads

3 · The agent: an orchestrator, not an inventor

A small agent on each machine configures workloads and reports health. Deliberately, it is deterministic code using proven tools — not an AI runtime and not a homegrown algorithm engine, because an operator has to be able to explain its behaviour at 3 a.m.

  • Always on: health reporting, heartbeat to the control plane, service discovery, workload-profile parsing.
  • Toggleable: workload-specific cache configuration, backup orchestration, restarting failed services, network monitoring.
  • Explicitly not in scope: custom caching algorithms, predictive prefetching, cache federation, model-based tuning, network-level optimisation (that belongs to the kernel and the tunnel, not the agent).

4 · Sovereignty: one interface over every provider

Providers have different APIs. Hardcode one and the architecture is locked to it. All machine and volume operations go through a small provider-abstraction layer, so no provisioning script contains provider-specific code:

# The shape of the abstraction (illustrative)
create_volume() {
  case "$CLOUD_PROVIDER" in
    digitalocean) doctl compute volume create "$1" --size "$2" --region "$3" ;;
    aws)          aws ec2 create-volume --size "$2" --region "$3" --volume-type gp3 ;;
    gcp)          gcloud compute disks create "$1" --size="$2GB" --zone="$3" ;;
  esac
}

The rule is simple to state and easy to break: no provider-specific call outside this layer. Sovereignty here is not a policy statement; it is the property that you can leave.

5 · The HTTP-only boundary: the edge never sees storage

The edge layer never knows block storage exists. Edge and node communicate over HTTP through a secure tunnel:

Edge worker                        Stateful node
    │  HTTP: POST /api/db/query         │
    │ ─────────────────────────────────►│  internally: local SSD cache,
    │                                   │  network volume, containers
    │  HTTP: 200 OK { results: […] }    │
    │ ◄─────────────────────────────────│
    │  (the worker knows nothing        │
    │   about storage)                  │

Separation of concerns, in one boundary: the edge is stateless and global, the node is stateful and regional, and neither knows the other's internals. Cross that boundary with a storage detail and you have started rebuilding the cloud you left.

Where AI Fits — and Where It Must Not

This architecture is deliberately unexciting about AI, and that is the position worth stating in a collection about AI-assisted work.

AI accelerates the building, not the running. The agent's runtime is deterministic on purpose; generated code is held to the same review and gate as anything else. The failure mode to avoid is an agent whose behaviour at 3 a.m. no operator can explain — an "inventor" where an orchestrator was needed.

The collection's thesis is load-bearing here. AI amplifies whatever structure it is given: inside clear boundaries it produces more of that clarity, and inside a tangle it produces more tangle. A sovereign fleet is what the fundamentals look like when the stakes are physical — a machine that fails, a volume that detaches, a restore that either works or does not. And the honest limits below are the part that makes delegating any of it defensible.

What This Is, and What It Is Not

This is: a reference architecture for stateful workloads on sovereign infrastructure; single-node per workload (not HA, not distributed); a recoverable pet (5–30 minute recovery); regional, because volumes are zone-bound; appropriate where 99.9% availability is enough.

This is not: a replacement for a managed database when you need a 99.99% SLA; multi-node high availability (no clustering, no leader election); zero-RPO (recovery point is the last backup, not the last transaction); serverless — you are managing the machine.

Use it when you need POSIX semantics, you want sovereignty from cloud vendors, you can tolerate a 5–30 minute recovery window, and you have the operational bandwidth to run infrastructure. Do not use it when any of those is false — the architecture is honest about that trade, and so should you be before adopting it.

The Cost Question

What a node costs is a fair question, and it deserves a precise answer rather than a confident one.

ComponentPlanning figure
Droplet, 1 vCPU / 1 GB (Sydney)$6 / month
Block volume, 100 GB$10 / month
Object backup storage (~20 GB compressed)$0.30 / month
Object backup egress$0 (zero-egress design)
Marginal cost per machine~$16.30 / month

These are the author's own working figures — a planning model, not an audited invoice. They are marginal cost per node, and the total cost of ownership is larger: the control plane must run somewhere, the node's egress to the edge is not included, operational time is the real cost, and system-disk and configuration backups are separate. What is not measured yet: p95 latency under production-shaped load, per-node support labour, egress and transit, and the hardware replacement rate. Those are the lines that would turn the model into evidence, and none of them is going to be invented here.

The Open Problems

Two parts of this architecture are design, not implementation, and it would be dishonest to present them as finished.

Database backup consistency. A file-level snapshot of a live database can capture an inconsistent state. The options under evaluation are the database's own continuous archiving plus a snapshot filesystem — the shape a PostgreSQL workload would use: base backup plus write-ahead-log archiving for point-in-time recovery, a filesystem snapshot for generic block volumes. Target RPO is one hour with continuous archiving (24 with nightly snapshots); target RTO is five minutes for a volume reattach, up to thirty for a full restore.

The security model. The node may run untrusted customer workloads, so compromise must be contained to that node — no lateral movement to the control plane, no access to other nodes' volumes. The layers are defined and the controls are not yet chosen: encryption at rest with its key management, how the node receives database credentials, how the agent authenticates to the control plane, container isolation level, and least-privilege provider tokens. A security review is a precondition of production use, not a follow-up.

The Orchestrator's Takeaway

Run stateful workloads where you control the exit — data on a detachable volume, the system disk disposable, the cache in front of the volume, and the edge kept ignorant of storage. Then say what it really is: a recoverable pet with a 5–30 minute recovery window, not cattle. The honesty is what makes it safe to rely on.

  1. The Second Life of Computers — the case for the fleet, and the pattern this node completes.
  2. The Edge-Native Stack — what the edge layer relocates back into the application.
  3. The Sovereign Fleet series (in preparation) documents the build layer by layer, and cites this architecture rather than repeating it.

For the full method, see The DevOps Engineer's Guide to Effective AI Usage.


Platform Engineering
Infrastructure
Sovereignty & Sustainability