lounge.

The Tree and the Flow

The idea

The city analogy

A large software system is hard to understand, and the difficulty has a specific shape. The flow — the sequence of steps that actually happens when something arrives — is nowhere written down. It is spread across the code, because code is conventionally organised by technological affinity: interface code here, API code there, business logic somewhere else. Functions from three layers run one after another in a single thread with no relationship to where they sit in the repository. Event-driven systems are harder still, since the order of execution depends on what arrives at runtime.

The flow therefore exists only at runtime and cannot be inspected directly.

A modern city offers a useful comparison. Its many independent institutions and millions of inhabitants make it complex, yet a first-time visitor can still navigate it.

Cities manage it because they are built around an organizing principle — highways, roads, streets. That infrastructure is not documentation added afterwards; it is what the city is made of. And because it is there, two things become possible that could not otherwise exist:

   an organizing principle   ──▶   an abstraction of it   ──▶   automation
   (roads and streets)             (the street map)            (turn-by-turn)

A street map depends on streets, and automated directions depend on the map.

Notice also what the streets are not: the inside of anything. Each of the city's institutions — schools, factories, hospitals, grocery stores, companies — has its own internal complexity. But the choreography that keeps city life flowing is shared, and it is separate from every one of them.

Large software systems often have no comparable organizing principle. No layer of the system is responsible for stating where things go, so there is nothing to abstract into a map and nothing for such a map to automate. The result is a steeper learning curve and more time spent reconstructing the flow.

The tree

The streets, in Lounge, are a tree of steps: routing nodes that decide where work goes, and task nodes that do the work. Deciding is separated from doing, deliberately and completely — the flow of events lives in the routing structure, the processing lives in ordinary functions, and neither is tangled in the other.

That separation makes the flow explicit rather than emergent. You describe it in one project.yaml, which serves as a table of contents for the application: the code supplies the chapters and the platform binds them.

The rest of the analogy follows in order:

  • The organizing principle is the tree, and the engine routes events through the very structure you described — not through a translation of it, and not alongside a picture of it. So the map cannot drift from the territory.

  • The abstraction is a real map you can open: a browser view that renders the application, its nodes, and its live traffic. Here is one, generated from a project.yaml and nothing else:

    The Functional View of a small application, rendered by the docs portal

    Nothing was drawn by hand. The portal read the spec, and this is what it found: an entry point that branches by input type, an ingest path that validates before it persists, an alert path that enriches and then fans out to three notifications at once, and a handler standing by for anything that fails. The colours are node types, so a glance tells you where the decisions are, where the work is, and where the one query sits.

    A codebase organised only by layers cannot produce that picture because its flow is not stated explicitly. Here the flow was written down first, so the picture is a reading of the program rather than a separate description of it.

  • The automation is everything the platform can now do on your behalf, precisely because it knows the shape of your system.

The organizing principle has two consequences that can already be checked.

A structure can be inspected; scattered behaviour cannot. A tree is a thing. It can be read, drawn, and — more usefully — checked. Lounge refuses to build a tree whose shape does not make sense, before a single event has run.

The engine decides what happens next, not your code. Your functions take a value and return a value. Everything between them belongs to the platform, because the platform is what walks the structure.

The rest of this wiki examines the consequences of that second property:

  • Two events arrive for the same customer at once. What happens?
  • A step produces something a later step needs. Where does that live in the meantime?
  • The work finishes. What goes back to whoever asked?
  • The process restarts. What survives?

The third is answered before this page ends. The others have pages of their own, arranged so that each answer builds on the preceding one.

Notice what those four questions usually are, though: ordering, state, the shape of your API, and recovery are normally four separate subsystems, with four vocabularies, none of them aware of the others. The claim this wiki has to earn is that here they are four consequences of one structure — which is the same reason a city needs one road network rather than four.

The remaining pages supply the mechanisms and evidence for that claim.

Prior art

Describing a system and having something run the description is an old and well-populated idea. Executable models, workflow and orchestration engines, state-machine services, flow-based programming tools — the field is decades deep, and Lounge claims none of it.

What is specific here is narrower, and it is the thing the rest of this page is about: routing is separated from processing. The tree carries the flow; your functions carry the work and stay ordinary Java, with no framework types in their signatures and no base classes to extend. The structure is not a notation you compile into a program — it is what the engine walks at runtime, which is why the same description can also be the thing you read, visualise, and reason about.

The rest of the wiki provides the basis for judging whether that separation is useful.

There is one less visible consequence. Because every application on the platform has the same shape, that machinery can be built once, into the scaffold, and inherited by everything standing on it. You do not assemble the parts and wire them together; you describe your tree, and they are already underneath it. That is what "no glue layer" means here.

An application built on Lounge is, naturally, a lounge — a place where your entities (orders, devices, players, conversations) each get a seat, their events are handled strictly in the order they arrive, and thousands of them are served at once without interfering with each other.

A few names, sorted out before they start appearing everywhere:

Name What it is
Lounge The platform and its engine — what you build on, what you deploy
LoungeTech The company behind it
AppTree The tree model the engine executes — what your project.yaml describes
JEQL Java Events Query Language — the query language that runs inside your flows

In these terms, you write an AppTree and Lounge runs it.

(Footnote for code spelunkers: config keys, packages and file magics all read lounge.*. Same engine, same names.)

The engine on one screen

Before examining the walk, here is a compact map. Each part receives its own page later.

The tree is the application. Routing nodes decide where an event goes; task nodes do the work — your Java function, a query, a SQL integration, a validator. Events enter at the root, route downward, and results bubble back up through the same topology.

Events belong to groups. A key function decides which entity an event belongs to. Events in a group are handled in order; different groups never wait on each other, and thousands run concurrently on virtual threads. The group is also the unit of isolation and, in a cluster, of placement. → Event Groups, Ordering and Parallelism

The root is a membrane. What escapes the tree unrouted is the answer. An envelope mode says how much of it you want back — the whole session, just the answer, or a bare acknowledgement — and the same contract holds over HTTP and over the native wire.

Records accumulate; queries read them. A task that returns a record has stored it, without an insert. Queries run inside the flow over records committed earlier, with grouping, projections and indexing. No ORM, and no external store for the working set. → Records and Queries

Durability is a log, not a store. Records live in memory; committed records and completed sessions are appended to a commit log so a restarted process replays warm. Queries never read the log — the record spaces rebuild from it. → Durability, Clustering and HA

Concurrency has three independent axes: unlimited between groups, an optional overlap within a group, and an optional fan-out within a single flow.

Distribution is topology. A session-affinity router places groups on workers and survives worker churn; internal edges can ride a framed binary wire on their own port; subtrees can live in other processes.

Operations are part of the product — severity-split logging with archival, provider-neutral OIDC/JWT authentication, health probes, ops surfaces, Helm charts, live statistics, and a browser UI that visualises the running project. → Building, Operating and Securing

The workflow

  1. Write plain Java records for what arrives and for what each step produces.
  2. Describe the tree in project.yaml — routing nodes, task nodes, queries, storage decisions.
  3. Run the linter; fix what it refuses.
  4. POST events at it — the tree carries them through your code.

One artifact serves every edition, with the same spec language and the same answers; what you pay for is speed, durability and topology, never different behaviour. Editions lists what each edition includes.

The walk

The walk is the basis for the mechanisms described in the rest of the wiki.

Every Lounge application is a tree. Not a tree of classes or of modules, but a tree of steps — the path work takes through your system, written down. An event enters at the root, walks the branches, and comes back out. That walk is the program.

The rest of this page explains how the walk works: what the nodes are, how an event moves between them, what happens when something fails, and how an answer finds its way back out.

The two kinds of node

There are two kinds of node.

A task node does something. It calls a function you wrote, or an integration that talks to the outside world. It takes one value in — and hands back one value, a sequence of values, or nothing at all.

The function is ordinary Java, from java.util.function. A plain task is a Function<In, Out>:

public static Contact normalize(ContactInput in) { ... }   // Function

No base class, no annotations, no engine types in the signature. Your business logic does not import Lounge, and it can be unit-tested with a plain call.

A task that needs configuration takes it as a first parameter (BiFunction<Config, In, Out>); that is how the built-in integrations are written, and the Integrations page is about them.

The middle case of the return contract deserves a closer look. When a task returns a collection or an array, the engine does not pass the collection along as a single payload. It emits one event per element, each routed independently from that node's position, each stamped with its index in the sequence. One task can therefore turn an order into three shipments, and each shipment continues down the tree on its own.

                              ┌──> Shipment 0   (index 0)
   Order ──> split-shipments ─┼──> Shipment 1   (index 1)
              (one task)      └──> Shipment 2   (index 2)

   one value in, three events out — each routed on its own

Returning nothing is the third case, and it is not a null. A task that returns null has broken its contract, and the engine says so. Absence is a value like any other: return an Optional, and let a branch decide what an empty one means.

The specialised task nodes narrow this contract deliberately: a validator returns a verdict, never a sequence; a filter passes one value or none; a consumer ends the line. Only the general-purpose task node fans out.

A routing node decides where an event goes next. It does no work of its own. Four routing shapes cover nearly everything:

Node Behaviour
SequentialRoutingNode Its children run in order, one after another.
BranchToFirstMatchNode The first matching child wins; its siblings are disqualified.
ParallelRoutingNode Every matching child runs. This is the fan-out node.
LoopRoutingNode Its two children repeat: work produces, a predicate judges, iterate prepares the next round.

These four shapes form the routing vocabulary. A pipeline nests them, and that nesting defines the control flow.

   event in
      │
      v
   root                          SequentialRoutingNode
      │
      v
   main-flow                     BranchToFirstMatchNode
      ├── raw.Order  ──> order-flow      SequentialRoutingNode
      │                    ├── normalize          TaskNode
      │                    ├── validate           ValidatorTaskNode
      │                    ├── enrich             TaskNode
      │                    └── notify             ParallelRoutingNode
      │                          ├── customer     TaskNode
      │                          └── warehouse    TaskNode
      │
      └── raw.Return ──> return-flow     SequentialRoutingNode
                           └── ...

The picture states the application's order of operations; there is no separate flow hidden in the implementation.

The sequential node

Of the four, you will reach for SequentialRoutingNode most, because it says the most ordinary thing there is to say about work: these steps, in this order. It is a process, written down.

The other three are the exceptions you reach for when a process stops being a straight line — when it must choose a path (BranchToFirstMatchNode), genuinely do several things at once (ParallelRoutingNode), or try again until the result is good enough (LoopRoutingNode). Most of a real application is none of these; it is one thing after another.

And because a sequential node's children may themselves be sequential nodes, a step can expand into a process of its own:

   fulfil-order                    ← a process
     ├── validate
     ├── reserve-stock             ← a step …
     ├── take-payment              ← … that is itself a process
     │     ├── authorise
     │     ├── capture
     │     └── record-ledger-entry
     └── notify

This is the ordinary act of breaking a task into sub-tasks, with the tree recording the result. "Fulfil an order" becomes four steps; "take payment" turns out to be three more; each of those could open further if it earns it. You decompose exactly as you would on paper, and the decomposition is the program.

The practical consequence is a style rule worth adopting early: name a sub-process for what it accomplishes, not how it works — take-payment, not payment-helper-chain. Done consistently, the tree reads as a description of the business, and someone who has never seen the code can follow it.

Branches and leaves

The two kinds occupy fixed positions:

Routing nodes are the internal nodes of the tree. Task nodes are always leaves.

A routing node has children and no work of its own; a task node has work and no children. Every path from the root therefore reads the same way — some decisions, then one thing done.

That shape is what makes routing tractable, because finding the next step is a recursive descent. A routing node asks its own policy which children to consider — the next sibling, the first match, all of them — and then asks each of those children the very same question. The recursion walks down through nested routing nodes and bottoms out at task nodes, which answer only for themselves: do I accept this value? A task accepts if the payload matches its declared input type and passes whatever subtype, predicate, and role checks it carries.

What comes back is a flat set of accepting leaves — the tasks that will actually run.

   SequentialRoutingNode
        │ asks
        v
   ParallelRoutingNode
        │ asks each child
        ├──> notify-customer    accepts   ─┐
        ├──> notify-warehouse   accepts   ─┼──> the running set
        └──> audit              declines   ┘     (two leaves)
                                (wrong type)

Usually that set holds exactly one leaf. It holds more only when a ParallelRoutingNode sits somewhere on the path, because that is the one node that offers the event to all of its children at once — so several subtrees can accept the same event, and every accepting leaf runs.

Routing

Routing is stepwise rather than planned globally. This distinction accounts for many otherwise surprising routing results.

Every event is stamped with its origin node — the node that produced it. That stamp is the whole navigation system, and it works like a finger held in a book.

When a task finishes, it does not look for the next node. It hands its output up to its parent. The parent reads the stamp to find out where in its own list of children that event came from, and takes it from there: a sequential node routes to the child after that one. If the origin was its last child, the parent has nothing left to offer, so it hands the event up to its parent — and the same question is asked one level higher. That handing-up is what "bubbling" means.

   SequentialRoutingNode
      ├── child 0: normalize
      ├── child 1: validate
      └── child 2: enrich

   1.  normalize finishes, hands its output UP,
       stamped "from normalize"                     normalize ──┐
                                                                v
   2.  the parent reads the stamp: origin was          [ the parent ]
       child 0, so the next step is child 1                      │
                                                                 v
   3.  validate runs                                          validate

So the sequence in your YAML is not a script the engine executes top to bottom. It is a set of neighbours, and each step is decided fresh: given where this came from, what is next here?

Two things follow, and both are why the model is worth the adjustment:

  • A subtree needs no knowledge of what encloses it. It answers only about its own children. Lift it into another tree and it behaves the same — which is what makes composition and mounting possible at all.
  • An event that nobody accepts is not an error. It keeps bubbling. If it reaches the root still unrouted, that is not a dead end either; as we will see, it is the answer.

Lineage

The stamp does more than navigate. An event a task produces is minted from the event the task was given: it joins the same group and the same session, and it points back to its predecessor by index. Counters travel the same way — a retry count, a loop count. Follow the chain of predecessors back and you reach an event with no predecessor at all: the external arrival that started the flow.

That chain is how the engine knows, at any node, which arrival a value descends from, who sent it, and with what roles — and none of it passes through your functions. A task sees a value and returns a value; the lineage is kept beside it, by the engine, for as long as the session lives.

The one place an event is not simply re-minted is a scope boundary. There it is wrapped, the way a letter goes into a new envelope: the group membership changes, the lineage inside does not. Subtrees, Mounts and Scopes is where that matters.

Unaccepted events

The recursive descent can come back empty: no leaf under this node accepts the value. This has a specific consequence:

A type mismatch between two nodes does not throw. The event simply stops being routed and bubbles up.

This produces silence rather than an error. The behavior lets a branch decline an event so a sibling can take it, but it also means a genuine mistake can look like nothing happening. The defence is not a runtime check; it is the linter, which reads your declared types at build time and refuses a tree whose steps cannot connect. Treat a lint finding as a failing test, because that is exactly what it is.

Branches discriminate two ways:

  • By type — a child whose first task takes domain.Refund matches only refunds.
  • By data — a routingPredicateRef inspects the value itself, so "orders above ten thousand" can take a different branch from ordinary ones.

The two combine: a child can be typed and predicate-gated, and both must pass before it accepts.

Branch and parallel matching

The two multi-child nodes differ in what they do with the leaves the descent turns up.

BranchToFirstMatchNode gathers the accepting leaves across all its children, then discards every one that does not sit under the first child that matched. It is a choice, not a broadcast.

ParallelRoutingNode keeps them all. Every accepting leaf runs. By default they run sequentially in child order on the calling thread — the "parallel" in the name is about semantics (all branches happen), not threads. Actual concurrency is opt-in, and it is the subject of the next page.

Bounded iteration

Some operations require several passes: for example, revising a draft until it reaches a quality threshold. LoopRoutingNode represents this explicitly with two children. The work child produces a candidate result; the iterate child converts a rejected result into the input for the next pass.

A predicate evaluates each work result using while-loop semantics: true requests another pass and false completes the loop. On completion, the work result bubbles up as the loop's output. Otherwise, the result passes through the iterate child and its output becomes the next input to work.

Two settings are mandatory, and the build refuses a loop missing either:

- type: LoopRoutingNode
  id: refine
  predicateRef: app::keepGoing     # true = another round
  config:
    maxCount: 5                    # finite upper bound
    policy: best-effort            # or: best-effort-not-enough
  children:
    - # work — produces the next result
    - # iterate — prepares the next attempt

maxCount provides the required upper bound. policy defines what happens when that bound is reached. With best-effort, the last work result is returned. With best-effort-not-enough, the loop raises a LoopExhaustedException. The exception carries the final result so an exception handler can report it or recover from it.

An exception thrown by either child ends the loop through the normal failure path. The iterate child handles results that need another pass; it does not handle failures. The round count travels on the event itself, as part of its lineage, so concurrent traversals keep independent loop state.

Keep the predicate inexpensive and observational. If evaluation requires a scoring pass or validation suite, put that step at the end of the work subtree and include its result in the output record. The predicate can read the score, and the iterate child can use the same value when preparing the next input. This avoids computing the assessment twice and retains it in the event history for later analysis.

Failures

A thrown exception does not unwind the tree. It becomes an event, of type Throwable, and it routes like any other event — looking for a node that accepts it.

That node is ExceptionHandlerTaskNode, and where you place it is a design decision, because placement determines reach: it handles failures that bubble up to it. Put one arm at the root and it catches everything the tree throws; put one inside a subtree and that subtree handles its own trouble.

   root
    └── lanes                         BranchToFirstMatchNode
         ├── the working lanes
         │        │
         │        └── throws ──> the Throwable becomes an event
         │                                  │
         └── on-exception <─────────────────┘   ExceptionHandlerTaskNode
                  │
                  v
           SanitizedException  ──> bubbles to the root and out

The handler logs the failure and replaces the raw Throwable with a small SanitizedException record — the type, the message, a bounded stack — because a live exception retained in a long-running session leaks memory.

A handler can do three more things, each an option in its config. It can retry, re-entering the failed node after a delay — retries bypass normal ingress, so design a retried step to be idempotent. It can close the group (closeGroupOnError: true): a one-way transition after which every later event for that key is refused, for the cases where continuing would compound the damage. And it can narrow what it catches, by exception type or by predicate, so a payment failure and a validation failure meet different arms. The handler reference has every option.

An exception arm is not optional decoration. Without one, a failure bubbles past the root and the caller receives an empty answer — which reads as "no result" when the truth was "it broke."

The membrane

The root node is a membrane, and the engine takes that literally in both directions.

Inbound, a record crosses downward exactly once: the request body is deserialised straight into your entry record in a single pass, and the engine routes it down the tree.

Outbound is the interesting half. Task outputs bubble upward, node by node. A value that reaches the root with nowhere further to route has, by construction, crossed the membrane outward — and that is the flow's answer.

The return path requires no declaration. Going down, an event passes branch points, and a decision is taken at every one of them. Going up there is nothing to decide: a node has exactly one parent, so a result has exactly one way out. The output of the last task in a subtree bubbles to that subtree's root, and for the tree as a whole, that root is the membrane. The route back is not written anywhere because there is only one.

No annotation marks a response and no flag names a return value. The topology determines the response: what escapes the root is what the caller gets.

This is why an exception arm matters so much, and why an unmatched type is worth catching at build time. Both are the same question — what reaches the root? — asked from different directions.

A scope with a front door of its own is its own membrane, with the same rule evaluated at its own boundary; Subtrees, Mounts and Scopes shows it.

Envelope modes

How much of the flow the caller gets back is one header, Lounge-Envelope. returns hands back what escaped the root — the answer — and is what a request/response client wants. full hands back every event the session generated, for debugging. ack is a bare acknowledgement for fire-and-forget ingest. The default is returns; a deployment can change it (lounge.process.envelope), and the request header wins. Rejections and errors are always full, because a thin error body helps nobody.

Some rules to remember

Five things to carry away:

  1. Routing is stepwise. A node knows only its next step; events bubble when there is none.
  2. A type mismatch is silent. The event stops. The linter is your guard — run the strict-lint build gate and treat findings as failures.
  3. BranchToFirstMatch chooses; ParallelRouting broadcasts. Reach for the second only when you mean every branch. And remember the third way work multiplies: a task returning a sequence emits one event per element, no routing node required.
  4. Exceptions are events. Give the tree an arm, or failures leave as silence.
  5. What escapes the root is the answer. Responses are topology, not annotation.

Where to go next

  • Event Groups, Ordering and Parallelism — who runs at the same time as whom, and what "in order" guarantees.
  • Records and Queries — asking questions of what the flow has already recorded.
  • Deep reference: tools/lounge/references/runtime-semantics.md (the mental model in full), tools/lounge/references/exception-handler.md (every handler option), docs/architecture/membrane-flows.md (the membrane walked through, channel by channel).