lounge.

Building, Operating and Securing

The previous pages described the application itself. This page covers the work around it: checking the specification while it is written, observing the running system, and securing its entry points.

The material is divided into building, operating and securing.


1. Building

The linter

Lounge checks structural properties of the tree at build time and refuses a configuration that violates them. The checks cover the following known failure modes.

Check Refuses
record-flow a step whose output type the next step cannot accept
duplicate-id two nodes sharing an id
exception-reachability an executable node with no catch-all handler above it
dead-sibling a branch placed after a catch-all sibling, where nothing can reach it
query-under-parallel a record-reading query inside a fan-out
queried-never-persisted a query for a record type nothing in the tree produces
closed-context a deployment whose queries reach outside its own data island
mutable-storable-indexed a mutable stored type that a query indexes
routing-predicate-placement a routing predicate declared where nothing consumes it
scaling-placement a deployment whose scaling mode contradicts its topology
statistics-placement, scope-dispatch statistics and scope boundaries in impossible positions

These checks have one property in common: each is a question about structure, answerable without running anything. This follows from the explicit topology described on page one. A codebase organised by layers cannot be asked most of these questions at all, because the structure the question is about is not written down anywhere.

Two deserve special mention because they catch errors that are otherwise silent. record-flow catches the type mismatch that would make an event stop mid-pipeline without an error — the one failure mode in the engine that is genuinely quiet. And queried-never-persisted catches the query that returns nothing forever, which is indistinguishable from "no matches today" until you look very closely.

Treat a lint finding as a failing test.

The portal

Point the docs portal at a project and it renders the application: the functional view with its node graph, every node's function source, input and output records, queries and tests — plus whatever design documents the project carries.

You met its output on page one. Two uses beyond reading:

  • Onboarding. The generated map gives a new reader the declared flow and cannot drift independently of the spec.
  • Review. A structural change appears as a corresponding change in the rendered shape, making changes to the flow visible during review.

Testing

Testing proceeds in three layers:

  1. The build gate. Building the tree under strict lint is a test — the fastest one you own. Run it on every change.
  2. Plain unit tests. Your task functions are ordinary Java taking a value and returning a value, with no framework types in their signatures. They test with a plain call and no engine at all. This follows from keeping routing out of processing.
  3. End-to-end. Post an event, assert on what comes back out of the membrane.

2. Operating

Logs

Logging separates operational records from short-lived diagnostic output.

Warnings and errors go to a structured JSON file, rotated and archived — this is the record you keep, ship to object storage, and answer questions from later. Debug and above goes to a self-deleting file: verbose, useful for an hour, gone before it fills a disk.

The split avoids applying one retention policy to two different uses. The diagnostic stream may be verbose because it is short-lived; the operational record remains smaller because it contains only warnings and errors.

The Lounge CLI page covers archiving the kept stream.

Health and ops surfaces

Health probes are independently switchable, because "is this process alive" is a different question from "may you ask this process anything else."

Everything else operational — the ops endpoints, statistics, cluster administration — lives behind a surface set you choose:

lounge.ops.surfaces=rest     # rest | none | your own

none is a real choice, and the reason it exists is worth stating: in a locked-down deployment the right number of administrative endpoints may be zero. You can also supply your own surface implementation, which is how an operator with an existing control plane exposes the same information through their own channel instead of a second one.

Statistics

Windowed statistics are declared as part of the tree, at a subtree root that becomes a statistics boundary. They appear as a node in the map rather than as separate metrics-library configuration.

Metrics can also be emitted mid-flow through the telemetry integration, which is the right tool when what you want to count is something only a step knows.

Record and replay

Because every completed session's external input is recoverable from the commit log, you can extract recorded traffic and re-inject it through a fresh tree against modified logic, then diff the outcome.

This permits a behavioral comparison against recorded inputs when asking whether a change alters the outcome.

Unlike crash recovery, which restores state without executing anything, this path runs the flows and their integrations. It therefore belongs in a test environment with test credentials.


3. Securing

Authentication and roles

Authentication arrives through one of two mechanisms, both of which produce the same identity representation:

Mechanism Trigger
Bearer JWT Authorization: Bearer …, verified against your identity provider
Form session a browser cookie

Both produce an identity with roles. Those roles reach the tree through one path and gate a requiredRole on a node:

- type: TaskNode
  id: issue-refund
  requiredRole: operator

That is deliberate. Authorisation is expressed in the same place as everything else about your application: the tree. A reader can see which steps are privileged by looking at the map, rather than by grepping for annotations.

The JWT knobs are provider-neutral — public key, issuer, audiences, and where in the token to find group claims. Any compliant provider works; none is special.

Protocols (cluster)

Roles say who may fire an event. A protocol template says when. A payment event, for example, may be valid in isolation but invalid before any item has been added.

A protocol gives an admission door's event groups a lifecycle grammar with three phases: one opening event, any amount of work, and one terminator:

protocols:
  point-of-sale:
    start: [domain.AddItem, domain.AddLoyaltyCard]
    work:  [domain.AddItem, 'domain.RemoveItem[admin]', domain.AddPayment]
    end:   [domain.Abort, domain.AddOrDeclinePhone]
    inactivityTimeout: PT15M

# bound where events enter — the root, or a scope with its own front door:
eventGroups: {store: local, groupBy: app::cartKey, protocol: point-of-sale}

Each group is traced through its phases: the first event must come from the start set and opens the group; work types repeat freely; an end type is admitted and then the group closes — nothing is admissible after. A role in brackets (RemoveItem[admin]) gates that one transition. Everything else — a payment first, an item after closing, a removal without the role — is refused at the door, before any node runs.

The refusal is deliberately terse toward the caller: a 409 naming the phase and a correlation id. The log line records the offending type, the group and the admissible set under the same correlation id. This limits information returned to a probing client while retaining enough detail for diagnosis. The violation also travels through the tree as an ordinary event, so an arm can count violations per caller, alert, or close the group.

Inactivity is also explicit. When a group remains idle past the protocol's inactivityTimeout, the engine admits a generic terminating event — ProtocolInactivityTimeout — through the same door, so your tree can react (release the reserved stock, void the transaction), and then the group closes. The timeout is part of the grammar, which is why declaring it is mandatory: every protocol says how long silence may last.

The contract survives restarts, and it is checked at build: a protocol naming a type the door cannot receive, or bound where there is no door, refuses before anything deploys.

Key errors

Authentication can otherwise fail silently: the process appears healthy while every secured request returns 500. Here a key file that does not resolve refuses at boot, and a remote issuer that cannot be reached is a readiness condition — an instance that cannot fetch the verification key stays out of rotation rather than accepting requests it cannot authenticate.

If the issuer URL is wrong, the instance never becomes ready, so the rollout stalls instead of completing. The deployment reports the problem before the instance serves customer traffic.

Defaults on

A set of protections that do not wait to be configured:

  • Response hardening — nosniff, frame denial, a restrictive content policy, and strict transport security when serving over TLS.
  • CORS closed by default — an allowlist you open deliberately, rather than a wildcard you remember to close.
  • Read deadlines — a client that opens a connection and dribbles bytes is timed out rather than allowed to hold a slot.
  • A connection ceiling, so exhaustion is bounded.
  • Redacted 5xx responses with a correlation id — the caller gets an identifier, the log gets the detail. Stack traces are not a public API.

The diagnostic information remains available through the correlation id without being returned to the caller.

Defaults off

Rate limiting and denial lists ship switched off, behind an interface you implement.

A generic rate limit can give a false assurance because it is not tuned to the application's traffic. The platform therefore provides the place to put the policy — evaluated at ingress, before work begins — while leaving the policy itself to the application operator.

Some rules to remember

  1. A lint finding is a failing test. The checks ask structural questions, which is only possible because the structure is written down.
  2. The map is derived from the specification, so it cannot drift independently and can serve as an onboarding document.
  3. Two log streams, two retention policies — verbose and disposable, important and kept.
  4. none is a valid surface set. Sometimes the right number of admin endpoints is zero.
  5. Authorisation lives in the tree, so privileged steps are visible on the map.
  6. An instance that cannot verify tokens takes no traffic. Local auth mistakes refuse at boot; a remote issuer is a readiness condition, so a wrong one stalls the rollout instead of serving 500s.
  7. Return a correlation id and log the detail, keeping diagnostic data out of the response.
  8. Rate-limit policy is application-specific. The platform provides the enforcement point; the operator provides the policy.

Where to go next

  • Durability, Clustering and HA — what survives a restart, and what each edition includes.
  • Deep reference: docs/ops/ — auth-modes.md, log-management.md, ops-surfaces.md, supabase-auth.md, clustering.md.