Building, Operating and Securing
The previous pages described the application itself. This page covers the work around it: checking the specification while it is written, observing the running system, and securing its entry points.
The material is divided into building, operating and securing.
1. Building
The linter
Lounge checks structural properties of the tree at build time and refuses a configuration that violates them. The checks cover the following known failure modes.
| Check | Refuses |
|---|---|
| record-flow | a step whose output type the next step cannot accept |
| duplicate-id | two nodes sharing an id |
| exception-reachability | an executable node with no catch-all handler above it |
| dead-sibling | a branch placed after a catch-all sibling, where nothing can reach it |
| query-under-parallel | a record-reading query inside a fan-out |
| queried-never-persisted | a query for a record type nothing in the tree produces |
| closed-context | a deployment whose queries reach outside its own data island |
| mutable-storable-indexed | a mutable stored type that a query indexes |
| routing-predicate-placement | a routing predicate declared where nothing consumes it |
| scaling-placement | a deployment whose scaling mode contradicts its topology |
| statistics-placement, scope-dispatch | statistics and scope boundaries in impossible positions |
These checks have one property in common: each is a question about structure, answerable without running anything. This follows from the explicit topology described on page one. A codebase organised by layers cannot be asked most of these questions at all, because the structure the question is about is not written down anywhere.
Two deserve special mention because they catch errors that are otherwise
silent. record-flow catches the type mismatch that would make an
event stop mid-pipeline without an error — the one failure mode in the engine
that is genuinely quiet. And queried-never-persisted catches the query that
returns nothing forever, which is indistinguishable from "no matches today"
until you look very closely.
Treat a lint finding as a failing test.
The portal
Point the docs portal at a project and it renders the application: the functional view with its node graph, every node's function source, input and output records, queries and tests — plus whatever design documents the project carries.
You met its output on page one. Two uses beyond reading:
- Onboarding. The generated map gives a new reader the declared flow and cannot drift independently of the spec.
- Review. A structural change appears as a corresponding change in the rendered shape, making changes to the flow visible during review.
Testing
Testing proceeds in three layers:
- The build gate. Building the tree under strict lint is a test — the fastest one you own. Run it on every change.
- Plain unit tests. Your task functions are ordinary Java taking a value and returning a value, with no framework types in their signatures. They test with a plain call and no engine at all. This follows from keeping routing out of processing.
- End-to-end. Post an event, assert on what comes back out of the membrane.
2. Operating
Logs
Logging separates operational records from short-lived diagnostic output.
Warnings and errors go to a structured JSON file, rotated and archived — this is the record you keep, ship to object storage, and answer questions from later. Debug and above goes to a self-deleting file: verbose, useful for an hour, gone before it fills a disk.
The split avoids applying one retention policy to two different uses. The diagnostic stream may be verbose because it is short-lived; the operational record remains smaller because it contains only warnings and errors.
The Lounge CLI page covers archiving the kept stream.
Health and ops surfaces
Health probes are independently switchable, because "is this process alive" is a different question from "may you ask this process anything else."
Everything else operational — the ops endpoints, statistics, cluster administration — lives behind a surface set you choose:
lounge.ops.surfaces=rest # rest | none | your own
none is a real choice, and the reason it exists is worth stating: in a
locked-down deployment the right number of administrative endpoints may be
zero. You can also supply your own surface implementation, which is how an
operator with an existing control plane exposes the same information through
their own channel instead of a second one.
Statistics
Windowed statistics are declared as part of the tree, at a subtree root that becomes a statistics boundary. They appear as a node in the map rather than as separate metrics-library configuration.
Metrics can also be emitted mid-flow through the telemetry integration, which is the right tool when what you want to count is something only a step knows.
Record and replay
Because every completed session's external input is recoverable from the commit log, you can extract recorded traffic and re-inject it through a fresh tree against modified logic, then diff the outcome.
This permits a behavioral comparison against recorded inputs when asking whether a change alters the outcome.
Unlike crash recovery, which restores state without executing anything, this path runs the flows and their integrations. It therefore belongs in a test environment with test credentials.
3. Securing
Authentication and roles
Authentication arrives through one of two mechanisms, both of which produce the same identity representation:
| Mechanism | Trigger |
|---|---|
| Bearer JWT | Authorization: Bearer …, verified against your identity provider |
| Form session | a browser cookie |
Both produce an identity with roles. Those roles reach the tree through one
path and gate a requiredRole on a node:
- type: TaskNode
id: issue-refund
requiredRole: operator
That is deliberate. Authorisation is expressed in the same place as everything else about your application: the tree. A reader can see which steps are privileged by looking at the map, rather than by grepping for annotations.
The JWT knobs are provider-neutral — public key, issuer, audiences, and where in the token to find group claims. Any compliant provider works; none is special.
Protocols (cluster)
Roles say who may fire an event. A protocol template says when. A payment event, for example, may be valid in isolation but invalid before any item has been added.
A protocol gives an admission door's event groups a lifecycle grammar with three phases: one opening event, any amount of work, and one terminator:
protocols:
point-of-sale:
start: [domain.AddItem, domain.AddLoyaltyCard]
work: [domain.AddItem, 'domain.RemoveItem[admin]', domain.AddPayment]
end: [domain.Abort, domain.AddOrDeclinePhone]
inactivityTimeout: PT15M
# bound where events enter — the root, or a scope with its own front door:
eventGroups: {store: local, groupBy: app::cartKey, protocol: point-of-sale}
Each group is traced through its phases: the first event must come from the
start set and opens the group; work types repeat freely; an end type is
admitted and then the group closes — nothing is admissible after. A role
in brackets (RemoveItem[admin]) gates that one transition. Everything else
— a payment first, an item after closing, a removal without the role — is
refused at the door, before any node runs.
The refusal is deliberately terse toward the caller: a 409 naming the phase and a correlation id. The log line records the offending type, the group and the admissible set under the same correlation id. This limits information returned to a probing client while retaining enough detail for diagnosis. The violation also travels through the tree as an ordinary event, so an arm can count violations per caller, alert, or close the group.
Inactivity is also explicit. When a group remains idle past the protocol's
inactivityTimeout, the engine admits a generic
terminating event — ProtocolInactivityTimeout — through the same door, so
your tree can react (release the reserved stock, void the transaction), and
then the group closes. The timeout is part of the grammar, which is why
declaring it is mandatory: every protocol says how long silence may last.
The contract survives restarts, and it is checked at build: a protocol naming a type the door cannot receive, or bound where there is no door, refuses before anything deploys.
Key errors
Authentication can otherwise fail silently: the process appears healthy while every secured request returns 500. Here a key file that does not resolve refuses at boot, and a remote issuer that cannot be reached is a readiness condition — an instance that cannot fetch the verification key stays out of rotation rather than accepting requests it cannot authenticate.
If the issuer URL is wrong, the instance never becomes ready, so the rollout stalls instead of completing. The deployment reports the problem before the instance serves customer traffic.
Defaults on
A set of protections that do not wait to be configured:
- Response hardening —
nosniff, frame denial, a restrictive content policy, and strict transport security when serving over TLS. - CORS closed by default — an allowlist you open deliberately, rather than a wildcard you remember to close.
- Read deadlines — a client that opens a connection and dribbles bytes is timed out rather than allowed to hold a slot.
- A connection ceiling, so exhaustion is bounded.
- Redacted 5xx responses with a correlation id — the caller gets an identifier, the log gets the detail. Stack traces are not a public API.
The diagnostic information remains available through the correlation id without being returned to the caller.
Defaults off
Rate limiting and denial lists ship switched off, behind an interface you implement.
A generic rate limit can give a false assurance because it is not tuned to the application's traffic. The platform therefore provides the place to put the policy — evaluated at ingress, before work begins — while leaving the policy itself to the application operator.
Some rules to remember
- A lint finding is a failing test. The checks ask structural questions, which is only possible because the structure is written down.
- The map is derived from the specification, so it cannot drift independently and can serve as an onboarding document.
- Two log streams, two retention policies — verbose and disposable, important and kept.
noneis a valid surface set. Sometimes the right number of admin endpoints is zero.- Authorisation lives in the tree, so privileged steps are visible on the map.
- An instance that cannot verify tokens takes no traffic. Local auth mistakes refuse at boot; a remote issuer is a readiness condition, so a wrong one stalls the rollout instead of serving 500s.
- Return a correlation id and log the detail, keeping diagnostic data out of the response.
- Rate-limit policy is application-specific. The platform provides the enforcement point; the operator provides the policy.
Where to go next
- Durability, Clustering and HA — what survives a restart, and what each edition includes.
- Deep reference:
docs/ops/—auth-modes.md,log-management.md,ops-surfaces.md,supabase-auth.md,clustering.md.