Event Groups, Ordering and Parallelism
The previous page followed one event through the tree. This page is about what happens when there are thousands of them at once — which of them wait for each other, which run side by side, and what "in order" actually promises.
Lounge uses the event group to answer those questions. It is both an ordering boundary and a storage boundary, so choosing its key is part of the application model rather than a scheduling detail.
The group
Every external event is asked one question on arrival: which group do you belong to? Your application answers it with a key function declared at the root: it reads the event and returns a key. The customer id. The order number. The channel name. Whatever your domain treats as "the thing this is about."
Events with the same key join the same group. Events with different keys can be processed independently.
Think of a motorway on which traffic cannot change lanes. Each group is a lane: traffic within it keeps its order and cannot overtake, while different lanes run independently. An event is assigned its lane on arrival, by its key, and stays there for its entire journey. More keys permit more independent work; ordering is preserved within each key.
Say the key is the order id. Then everything that happens to order 42 — its payment, its shipment, its invoice — belongs to one lane, however those events arrive interleaved with other orders':
arrivals, in the order they reach the door
──────────────────────────────────────────
t1 pay order-42
t2 pay order-77
t3 ship order-42
t4 refund order-77
t5 invoice order-42
the lanes they are assigned to, by order id
───────────────────────────────────────────
[ order-42 ] pay(t1) ──> ship(t3) ──> invoice(t5)
[ order-77 ] pay(t2) ──> refund(t4)
Inside a lane, that order is guaranteed: order 42 is never invoiced
before it ships. Between lanes, nothing is promised — refund(t4) may
well finish before ship(t3) does.
The resulting contract has two parts:
Within a group, events complete in the order they arrived. Across groups, there is no ordering at all — thousands of groups make progress at the same time.
This is comparable to the ordering obtained from a partitioned queue or a per-key lock. In Lounge it follows from the group key, without a separate queueing scheme in the application.
Group memory
A group is not only an ordering rule. It is the place records accumulate.
When a step commits a record, the record lands in the group's store — and a query in a later event of the same group can find it. That is why Records and Queries can talk about querying "what the flow already knows" without ever mentioning a database: the store is the group, and the group outlives the single event.
The corresponding design rule is that isolation is per group. Two customers with different keys cannot see each other's records, because they are not in the same lane. Choosing the session key is therefore a data-boundary decision, not just an ordering one.
Closing a group
A group has exactly two states: ACCEPTING and REJECTING. The transition
is one-way. Once closed, every later event for that key is refused — which
is what closeGroupOnError on an exception handler does, as the
previous page described. There is no reopening; a closed lane stays closed.
Ordering
The engine offers two options, distinguished by when an event waits for its predecessor.
Serial admission makes each event wait at the door. Event N+1 does not begin processing until event N has finished. The lane is strictly one-at-a-time.
Overlapped admission lets each event start immediately. It still waits for its predecessor — but later, and only when it must.
That later wait does not occur simply "at the end". An overlapping event runs freely until it needs to know something about the past, and a query is exactly that. When event N+1 reaches a query node, it blocks until event N has finished processing completely, because the answer must include everything N committed. If the event never queries, the wait happens at the exit instead, so that completion still lands in arrival order.
SERIAL e1 ├──── work ────┤
e2 ├──── work ────┤
waits at the door
OVERLAPPED e1 ├──── work ────┤
e2 ├─ work ─┤▓▓▓▓├─ query ─┤
└ blocked here: the query needs
everything e1 committed
So the rule is: overlapping buys you the work an event does before its first query. An event that queries as its first step behaves much like a serial one. An event that does substantial work and only then consults the record overlaps almost entirely. This is a design lever, not just a setting — placing a query later in a flow is what lets the overlap pay.
What the choice does and does not change:
- What changes: how much work inside a lane can proceed at the same time.
- What does not: the order in which events finish, or what they see. Both modes complete in arrival order, and a query always observes the complete history before it — which is precisely why the query has to wait.
The two modes produce identical answers and differ only in elapsed time, so enabling overlap does not require redesigning the flow.
Whether you need it is a question about your traffic, not your code: serial admission costs you only on a hot group doing real work per event, and nothing at all when a thousand groups each do one thing. Overlapped admission is a paid capability; without it the engine degrades to serial rather than refusing to run. The FAQ works the cost through.
Parallelism inside one event
Everything so far concerned different events. There is a second kind of concurrency entirely: one event that fans out inside the tree.
ParallelRoutingNode is where that happens. As the previous page noted,
it is a semantic fan-out first — every matching branch runs — and by
default those branches run sequentially, in child order, on the calling
thread.
Adding one line changes that:
- type: ParallelRoutingNode
id: notify
config:
execution: concurrent
Now each matched branch is forked onto a virtual thread and the node joins them before moving on. The observable contract remains the same: downstream sees no difference. Events bubbling up from the forked branches are held at the join and replayed sequentially in child order, so everything after the node observes exactly what sequential mode would have produced. The option changes elapsed time, not the ordering visible beyond the node.
Branch failures
Fanning out raises a question sequential code never has to answer: one
branch failed — what about the others? By default the failure rides onward
as an exception event and the successful siblings survive beside it, which
is what independent branches want: one notification failing should not
cancel the rest. Two stricter, all-or-nothing policies exist for work that
is pointless once any part fails (failure: fail-fast, failure: collect-all); the parallel-routing reference describes both.
The fold
Fan-out has a counterpart. A reducer folds the branches' outputs into a single value at the join:
config:
reduce: first # or reduceFnRef: app::pickBest
first keeps child 0's answer; a reduceFnRef folds the results pairwise
in declared child order. Without a reducer the node replays every output —
fan-out. With one, it produces exactly one — fan-in. The reducer only sees
branches that succeeded.
The fold does not need execution: concurrent: which of three quotes to
keep is a question about meaning, not threads, so the node folds the same
way on every edition and concurrency only makes it faster. Fan out to ask
several sources, fold to one answer — that is map-reduce, expressed as
topology rather than as a library call.
Queries inside a fan-out
A record-reading query cannot run inside a fan-out. The reason is the same one
that governs admission: a query is a synchronisation point and must not run while
something that could write the records it is looking for is still running.
Between events, the engine enforces that by waiting. Inside a fan-out it
cannot, because the branches are peers — nobody is "earlier" than anybody
else. So a FIND in one branch would see whatever its siblings happened to
have finished.
The linter therefore rejects a query node with a parallel-routing node
anywhere above it in both execution modes. Sequential branches happen to
produce a deterministic answer, but adding execution: concurrent later — or
deploying the same tree on an edition that has it — would change that answer.
The linter rejects a design whose query semantics depend on the threading mode.
The shape that is correct puts the query after the fan-out, where the branches have all finished by construction:
parallel: ask-each-division <- branches write records
├── division-a
├── division-b
└── division-c
query: FIND … GROUP BY division <- sees all three
That is the ordinary scatter-gather, and it works precisely because the sequential parent will not offer the query node the event until the fan-out step is done.
One thing the rule does not forbid, because "query" is doing two jobs in
that sentence: a projection. A FIND searches the record store; FOR q IN input PROJECT INTO Quote { … } reshapes the value it was handed and reads
nothing. The linter knows the difference the same way the node does — a
projection compiles to an implementation with no access to a store — so
projections are welcome inside a branch, which is convenient, since giving
each branch's result a common shape before folding is what a map step is
for:
parallel: ask-each-source reduce: cheapest
├── source-a → PROJECT → Quote
├── source-b → PROJECT → Quote <- projections: fine
└── source-c → PROJECT → Quote
Cross-group views
If no query crosses a group, how does anyone see a total per department, or across the whole system? A subtree root can declare a new grouping, so the same events are seen per session in one part of the tree and per department, or globally, in another. Subtrees, Mounts and Scopes shows how, and what the global case costs.
Some rules to remember
Five things to carry away:
- The group key defines the lane. It sets both the ordering boundary and the isolation boundary — choose it as a data decision.
- In-group order, cross-group freedom. Events in a group complete in arrival order; different groups never wait on each other.
- The two admission modes differ only in when an event waits — at the door, or at its first query. Same answers, same completion order, different wall clock.
- Concurrent fan-out changes latency, not semantics, and the fold works in either mode — map-reduce is expressible on every edition.
- No searching inside a fan-out. A
FINDgoes after the branches, not in one; a projection, which searches nothing, is free to sit inside.
Where to go next
- Records and Queries — querying the records a group has accumulated, including from inside the same flow that wrote them.
- Subtrees, Mounts and Scopes — where a subtree chooses whose records its queries read, and how re-grouping works.
- Deep reference:
tools/lounge/references/runtime-semantics.md(admission, fan-out and query semantics together), and the design records for concurrent routing and map-reduce underdocs/superpowers/specs/.