lounge.

Durability, Clustering and HA

Everything so far has described one process holding its state in memory. This page is about what happens when that is not enough — when the process restarts, when one machine is not big enough, and when a machine disappears.

The three mechanisms depend on the same separation. Durable artifacts — log segments, snapshots and manifests — are immutable and portable; the processes that produce them are replaceable.

Durability and the commit log

Records are held in memory, which permits queries to run efficiently inside a flow. Durability is handled separately rather than by making the record spaces themselves persistent.

Every committed record and every completed session appends to a durability log. It is write-behind: nothing reads from it during normal operation. Queries never touch it. The log exists to reconstruct state after a process restart.

   in memory:   records ── queries read these ──▶ answers
        │
        └─ appended ──▶  segment log  ──▶  read only when a process restarts

When durability is disabled, the append path is absent; there is no dormant log operation on the normal processing path.

There are two durability levels, distinguished by what may be lost. async batches the force into the background, with low overhead and a bounded loss window. msync makes a session's completion wait for the forced watermark, so no completed session is ever lost; its throughput is consequently bounded by the storage device's fsync rate. The choice is between that completion guarantee and the bounded loss window of asynchronous forcing.

File and database backends

Commit-log entries can use either of two backends. Both implement the same contract, recovery process and licensed capability, so the choice depends on the deployment environment.

A file commit log is self-contained and has the highest throughput: the process owns its storage and has no external dependency. It requires a volume that outlives the process and can be re-attached under a stable identity. On a container platform, that requirement must be reflected in the storage design.

A database puts the same entries in a table. Where the deployment already has a managed, backed-up, replicated and monitored database, the commit log can use those facilities without requiring a separate persistent volume.

Both are switched on with the lounge.ha.log.* properties, and the backend changes nothing about recovery, the durability level or restart behaviour. The database backend batches its writes and needs a pooled connection — a handshake per batch gives most of the throughput away — and when the database is unreachable it holds the batch and blocks callers rather than dropping accepted data. docs/ops/ha-recovery.md has the keys and the measured figures.

File-log rotation bounds replay. When the log outgrows its threshold, live state is re-logged as a fresh generation and the old segments seal into an archive. The same process carries on over the new file. Replay time is thereby bounded by live state, not by how long the application has been running. A system that has been active for a year does not therefore require a year of history to restart.

Restart restores committed state. A restarted node reads its log and comes back with its committed state present: records go back into the record spaces and their indexes, and completed sessions are restored as readable history.

Recovery does not re-run flows. Nothing is routed through the tree, no task node executes, and no integration fires. Recovery rebuilds memory to match the log and then processes whatever arrives next.

A flow that charged a card, sent a confirmation and called a shipping API must not repeat those effects after every restart. For that reason, recovery replays state, not work.

(The same log also feeds the record-and-replay testing tool, which does run your flows; Building, Operating and Securing describes it.)

Subtree deployments

A subtree with its own store is, by construction, self-contained with respect to Lounge records. Nothing outside it feeds its queries, so it can run in a different process without changing what any of its queries see. Once it is running in its own process, with its own state and scaling policy, it has the operational properties normally associated with a microservice.

This boundary need not introduce a second application. Ordinarily a service boundary means its own repository, its own API contract, a client on one side and a server on the other, and two things to keep in step forever. Here the boundary is a few lines of the same description. The client and the server are derived from it.

Independent sizing and scaling

A separate deployment may also be sized separately.

Parts of an application can have different resource profiles. A subtree that waits on an external API is network-bound — it needs enough concurrency to hold many waits open, and almost no CPU to do it. A subtree doing real computation per event is the opposite. In one process they share a single resource envelope, so both are sized under one allocation.

Split, each deployment carries its own replica count and its own CPU allocation, scaled on its own signal. The network-bound subtree can use more low-CPU replicas, while the compute-bound subtree can use fewer replicas with larger CPU allocations.

Declaring it takes one block, and the scaling mode comes with it (the next section is about those modes):

deployments:
  - name: alerts-svc
    assignedNodes: [alert-flow]
    scaling: routed

alert-flow now runs in its own deployment. The spec does not change; the same project.yaml is deployed everywhere.

Stubs and proxies

At build time each process asks, node by node, is this mine?

    THE SPEC                  MAIN DEPLOYMENT           ALERTS-SVC
    ────────                  ───────────────           ──────────
    root                      root                      (not built)
     ├── ingest                ├── ingest               (not built)
     ├── alert-flow            ├── ◻ stub  ──────┐       ▣ proxy
     │    ├── enrich           │   (a stand-in)  └──▶     ├── enrich
     │    └── notify           │                          └── notify
     └── respond               └── respond               (not built)

Three outcomes, from the same file:

  • A node assigned elsewhere, at a place your tree needs to route through, becomes a stub — a local stand-in sitting exactly where the subtree would have been. Your tree routes into it normally; it forwards the event to the deployment that owns it.
  • A node assigned to you and marked remotable gets a proxy at its root. The proxy accepts a forwarded event, reconstructs it, runs the real subtree, and gathers the outputs for the reply.
  • A node assigned to neither is simply not built. It does not exist in your process at all.

There is no separate "client" or "server" artifact to maintain. There is one description, and the shape each process takes is derived from which nodes it was assigned.

Delivery contracts

Delivery is exact by default: the stub waits for the subtree's outputs and they continue the flow. A delivery: lossy edge — fire and forget, failed forwards audited and dropped — exists for statistics and audit feeds where dropping under pressure beats back-pressure into the caller; the clustering reference covers it.

Storage for remote subtrees

A remoted subtree either owns its store, or has no queries. The enclosing tree's record space is in another process and is therefore unavailable to the remote subtree.

The closed-context rule checks this at build time. A subtree whose queries read only what it wrote behaves the same in-process and out, so an invalid dependency is reported before deployment rather than later through an empty result set.

Clustering by event-group ownership

Affinity

A replica holds session state in memory. With two replicas behind an ordinary load balancer, two events from the same session can land on different machines, each holding only part of the session state.

By contrast, a request-scoped flow backed by SQL does not care which replica serves it. The problem is specific to state the engine keeps in memory on behalf of the application.

The affinity router

The fix is affinity: every event group has exactly one owner replica. A process running with lounge.role=router sits in front of the workers, derives the group key from each incoming event, and forwards it to that group's owner.

                    ┌──────────┐
   event ──────────▶│  router  │  derive group key, look up its owner
                    └────┬─────┘
             ┌───────────┼───────────┐
             ▼           ▼           ▼
         worker A    worker B    worker C
        groups 1,4   groups 2    groups 3,5

An ordinary load balancer cannot provide this affinity because deriving the group key means running application code. The key comes from a function you wrote, over a payload only your types describe. A proxy that does not host your application cannot know which lane an event belongs to. So the router is an ordinary Lounge process in a different role, applying exactly the same derivation a worker would.

Router state and worker authority

The router's table of group-to-worker is derived from what the workers report, not a ledger the router owns. Workers are the authority on which sessions they hold; the router asks and caches. A router that restarts rebuilds its map by asking again.

Worker loss

When a worker is lost, its session state is lost with it. The router marks those groups rejecting and answers their events itself:

HTTP 409   { "status": "rejected", "reason": "worker-lost" }

Quietly placing the session on another worker would imply continuity over state that no longer exists, so the platform refuses to do it. The client would otherwise believe it was continuing a conversation whose memory had been lost.

After a timeout, the key may start a genuinely fresh session on a live worker; it is not treated as a resumed session. A brand-new group whose first forward hits a dead worker is re-placed because no session state existed on that worker.

Scaling modes

A deployment's scaling mode follows from what it holds, and the engine reads it off the topology: a deployment with no in-memory state is stateless and needs no router; one holding per-group state is routed — sessions have owners; one holding a global statistics boundary is singleton, because a group that aggregates all traffic cannot be divided by affinity. You can declare scaling: on a deployment to override the derivation, and the linter checks that a declaration does not contradict the topology.

Scale-down by draining

Automatic scale-down is deliberately disabled. Removing a replica that holds live sessions destroys them, and an external autoscaler does not know which sessions are still active.

Instead a worker is designated for retirement: it stops receiving new sessions immediately, and keeps serving the ones it has. Sessions are allowed to end rather than being moved. When the count reaches zero the worker reports itself retirable and can be removed with no race. Abandoned sessions are ended by the idle sweeper, which prevents an inactive session from retaining a worker indefinitely.

Internal edges — router to worker, subtree to subtree — can ride a framed binary wire on its own port instead of HTTP; it is optional per edge, HTTP stays the liveness path, and an edge that cannot establish the wire falls back to HTTP rather than failing. docs/ops/native-wire.md has the setup.

Instance rotation and identity handover

Instance rotation combines durability and clustering. A worker can hand its identity — its log, its sessions, and its place in the router's map — to a fresh replacement.

  1. The outgoing worker settles: it finishes what it is doing, rotates its log, and writes a manifest — a digest proving exactly which groups and how many rows its log carries. Then it seals, refusing new work and releasing its directory.
  2. The router holds that worker's incoming events. Acked work is journaled to a durable custody file; request/response work is parked, where the waiting client's own retry is the recovery.
  3. A replacement boots on the same log directory, replays it, recomputes the manifest digest, and advertises the result as an adoption proof.
  4. The router changes ownership only after verifying the adoption proof. The digests must match before the handover proceeds.

Before ownership changes, the handover can be aborted: the sealed worker can take its directory back and return to service.

The third and fourth steps prevent readiness from being inferred merely from a replacement process starting. The matching digest demonstrates that the replacement replayed the state described by the outgoing worker's manifest.

The Lounge CLI drives these operations — build, deploy, run, heartbeat and stop — and the Lounge UI invokes the same verbs through a local bridge. Both therefore use the same deployment and recovery path.

Operational constraints

  1. The log is a recovery artifact, not the query store. Nothing reads it in normal operation; record spaces rebuild from it.
  2. A file or a database — same contract. Pick by what your deployment already operates. If you pick the database, give it a pooled connection.
  3. Rotation bounds replay time by live state, not by uptime.
  4. Recovery restores state; it does not re-run work. No task node executes on replay, so nothing you did to the outside world happens twice. The record/replay tool is the opposite and belongs in a test environment.
  5. The group is the unit of placement, which is why the router must run your key function and cannot be a plain proxy.
  6. A lost session is reported rather than silently re-created. The router returns a 409 instead of implying continuity.
  7. Scale-down drains sessions. Existing sessions end on their current worker rather than moving.
  8. Identity handover requires proof. The router changes ownership only after the manifest and adoption digests match.
  9. A remoted subtree owns its store or has no queries. The closed-context check enforces it at build time, so a subtree behaves the same in-process and out.

What to read next

  • Editions — what each edition includes, how capabilities are enforced, and how billing works.
  • Operations detail: docs/ops/ha-recovery.md, docs/ops/clustering.md, docs/ops/native-wire.md
  • The compose drill, to run the handover yourself: tools/testing/handover-drill/
  • Operational tooling: Lounge CLI and tools/lounge/references/ui.md
  • Design records: docs/superpowers/specs/2026-07-24-ha-replication-design.md and the 2026-07-28 handover spec