Records and Queries
Every application eventually needs to remember something and then look it up. In most systems those are two separate exercises: a schema, a connection, a mapping layer, an insert on one side, and on the other a query language that knows nothing about your program.
Lounge treats the values already produced by the flow as the material for later queries. This page describes what is retained, where it lives, how queries read it, and how long it remains available.
Records
The storage rule is the useful place to begin: you never write an insert.
A task node returns a value. If that value is a record, the engine keeps it. The event continues to the next node while the record remains in the store, available to later queries.
event arrives ──▶ validate ──▶ price ──▶ reserve ──▶ confirm
│ │
▼ ▼
Quote Reservation <- kept, without
being asked
The store is therefore a record of values the tree has produced, rather than a destination to which application code writes explicitly.
Retention
The default retention rule has three parts: Lounge keeps what arrived from outside, what went wrong, and what your queries read. Other intermediate values are not retained.
External arrivals are marked when they enter the tree and survive whether or not a query refers to their type. Exceptions are also stored and queryable, regardless of the rest of the retention configuration. For the remaining records, the engine examines the queries in the tree and retains the types those queries read; those types constitute the application's working set.
There is one related refinement rule: a task whose input and output have the same type is treated as refining the input rather than adding another fact. Threading a record through five enrichment steps therefore leaves the final version rather than five nearly identical rows. This rule never retracts an external arrival, and a task that fails retracts nothing, so its input remains available.
Applications can override these defaults: a node may retain a plain object, treat a same-type output as a new fact, or select the debugging policy that retains everything. Those controls are described in the persistence reference.
Record shape
A task may pass any object — records, plain objects with public members, nested, collection-bearing. Storage, queries and the wire all accept them.
The principal hazard concerns mutability rather than shape. An index entry snapshots a value when the record arrives. Mutate that value afterwards and the index is quietly describing something that no longer exists. Immutable records make this impossible.
Prefer immutable value carriers for facts you intend to store or query. The machinery accepts mutable objects, but immutable values keep an index consistent with the value it describes. (A lint check covers the reachable case: a mutable, storable type that some query in the tree reads.)
Record stores
The store belongs to the event group — the lane from the previous page. Records therefore accumulate across every event in a group and remain invisible to every other group. The same boundary provides both ordering and isolation.
Records live in memory. That is a deliberate design choice rather than a limitation to work around, and the rest of this page is largely about its consequences: what you keep, how long, and where history goes when it outgrows the heap.
One more property, and it is the reason scatter-gather works at all: a query sees both the records committed by earlier events in its group and the ones the current event has produced but not yet committed. So a fan-out that writes a record per branch, followed by a query that rolls them up, is an ordinary pipeline rather than a separate request cycle:
parallel: ask-each-division ──▶ FIND … GROUP BY division
├── division-a → DivisionTotal sees all three, in this same event
├── division-b → DivisionTotal
└── division-c → DivisionTotal
Queries
In most event-driven architectures reading is kept apart from writing, known as CQRS, command-query responsibility segregation: the flow writes, a read model is built beside it, and a query service answers from that model. Here the tree is the flow, and a query is a step in it. A query node asks the group's records a question at the point where the answer is needed, sees everything committed before it, and hands the answer to the next step. There is no second model to keep in step, and a lookup, an enrichment or a rollup sits where it is used rather than in a service of its own.
JEQL is the language those nodes speak: one statement per node, three verbs,
each one aspect of the work. FIND searches the records for what you do not
hold. FOR binds a value you do hold, the input, the node's config, a found
record, and hangs nested searches off it. PROJECT reshapes the result into
a record you have declared. They are kept apart on purpose: a search does not
decide shape, a binding does not search, and a projection does not look
anything up.
The text of a query is the shape of its answer. A FIND hands on a list
of what it found. A FOR hands on a frame keyed by its alias and the names
of its nested searches. A PROJECT hands on the record it names. The node's
output type is read off the query, and the linter checks it against the next
node like any other seam.
That makes a query node the tree's type changer. It takes one type in and hands another out with no Java between them: a request becomes a request with its history attached, a list of sales becomes a regional total, a domain record becomes the summary an API wants. Most of the reshaping in a Lounge application happens in these nodes, declared rather than coded.
FIND and FOR
The two verbs divide the work as follows:
FINDsearches for records you do not hold.FORbinds a value you do.
FIND goes looking:
FIND o<Order> WHERE o.customerId = input.customerId ON_EMPTY EMPTY_LIST
FOR takes something already in your hand — the event payload, the node's
config — and gives it a name so you can hang other things off it:
FOR input<ShipmentRequest> {
orders: FIND eo<EnrichedOrderEvent> WHERE eo.customerId = input.customerId ON_EMPTY EMPTY_LIST
}
PROJECT INTO ShipmentDecision { shipmentId: input.shipmentId, orders: count(orders) }
Read FIND o<Order> as "find orders, and call each one o". The alias is not
decoration — it is how the WHERE clause names the fields of a record it has
not found yet, how a nested query refers back to this level, and, as you will
see below, how the answer itself is labelled. It is required.
This gives JEQL a consistent syntactic rule that the compiler can check:
Left is what you are looking for. Right is what you already have.
o.customerId = input.customerId reads "the order's customer id equals the one
on my input". Write it backwards and the compiler stops you. Every JEQL
condition has this shape: the sought record's field on the left, and on the
right a value you hold — input, config, an enclosing alias, or a literal.
Query results and cardinality
Two rules govern every result, and both are visible in the queries above.
A query returns a list. However many records match — none, one, or many —
the answer is delivered as one list. A single match produces a one-element
list, and no matches produces an empty list. Downstream, iterate it, reshape it
with FOR, or declare the next node's input as java.util.List and consume
the list directly.
An empty answer is the empty list, unless the query says otherwise. A
FIND may end with an ON_EMPTY clause, and needs one only when finding
nothing is a failure worth stopping for:
FIND o<Order> WHERE o.customerId = input.customerId
FIND o<Order> WHERE o.customerId = input.customerId ON_EMPTY RAISE_EXCEPTION
The first hands on [] when nothing matches; the second raises. A third word,
DELETE_PARENT, exists only inside nested searches and is described below.
Writing ON_EMPTY EMPTY_LIST is legal and says nothing the default did not.
One assertion is available on the other side of the count. If a query is
supposed to answer with at most one record — an identity lookup —
ON_MANY RAISE_EXCEPTION makes two or more matches an error. It asserts the
size of the answer; it does not make the records unique. Two committed copies
of the same fact trip it, and collapsing copies is DISTINCT's job:
FIND u<User> WHERE u.email = input.email ON_EMPTY RAISE_EXCEPTION ON_MANY RAISE_EXCEPTION
Thus cardinality has three explicit properties: the result is a list, the empty case is declared, and a single-record answer may be asserted when required.
The query's output
The query's output rule is particularly important when composing a pipeline.
A query node is a task node. Task nodes turn their input into their output, and
a query is no exception: what it hands to the next node is its answer, and
its answer is exactly what the query text declares. A bare FIND declares
the found records, so the found records are the payload and nothing else is:
input: AlertRequest ──▶ [ FIND InventoryUpdate … ] ──▶ payload: [ the matches ]
(AlertRequest: gone)
Most of the time that is exactly right: you asked a question, and the answer is
what you want to carry forward. When the next node needs both — the request
and what you found for it — the request has to be part of the declared answer,
and that is what FOR says. It binds the input under an alias, hangs the
results beside it, and then names what leaves:
FOR input<AlertRequest> {
lowStock: FIND iu<InventoryUpdate> WHERE iu.quantity < input.threshold ON_EMPTY EMPTY_LIST
}
PROJECT INTO LowStockContext { alertId: input.alertId, skus: lowStock.sku, found: count(lowStock) }
Inside the tail the frame is
{ input: <the AlertRequest>,
lowStock: [ <InventoryUpdate>, <InventoryUpdate>, … ] }
— the bound value under its alias, input here, and each nested search under
its declared name — and the answer is one LowStockContext. The frame is the
query's own structure; the record is what the next node sees. The text of the
query therefore describes both what it reads and what it hands on.
Nested queries
A FIND can carry a block of further searches, each correlated with the level
above it:
FIND c<Company> WHERE c.companyId = input.companyId ON_EMPTY EMPTY_LIST {
departments: FIND d<Department> WHERE d.companyId = c.companyId ON_EMPTY EMPTY_LIST {
employees: FIND e<Employee> WHERE e.deptId = d.deptId ON_EMPTY EMPTY_LIST
}
}
Correlation is by name, at any depth. The inner query says d.deptId
because d is in scope, and it could equally say c.region to reach two
levels up. There is no special syntax for "my parent" or "my ancestor" — naming
the alias is the whole mechanism, which means a nested query reads the way
nested scopes read in a program.
The shape of the answer mirrors the shape of the question: a list of
companies, each {c: …, departments: […]}, each department carrying its own
employees. Every nested field is always a list, empty included — the shape of
a result never depends on what the data happened to contain.
The third empty policy applies here. ON_EMPTY DELETE_PARENT, on a nested
search, means "if this inner search finds nothing, drop the enclosing
record from the results entirely" — the JEQL way of writing an inner join,
pruning recursively up the nesting. At the top level there is no enclosing
record to delete, and the compiler says so.
Projection
Records are the shape your tree produces. They are rarely the shape an API response wants. A projection closes that gap:
FOR o IN input
PROJECT INTO OrderSummary { id: o.orderId, total: o.amount }
Projections carry the vocabulary you would expect for summarising — LET,
WHERE, GROUP BY, HAVING, ORDER BY, DISTINCT, SKIP, TOP — and
they end in exactly one way: PROJECT INTO a record you have declared.
A FIND can end with the same clauses directly, which saves a node when all
you want is a summary:
FIND s<Sale> WHERE s.customerId = input.customerId ON_EMPTY EMPTY_LIST
GROUP BY region := s.region
PROJECT INTO RegionTotal { region: region, total: sum(s.amount) }
Fields are assigned by name, so the order you write them does not matter — the compiler reorders them into the record's canonical constructor and checks the set against its components. Get a field name wrong, or forget one, and the build fails. Declaring the record therefore makes the projected shape available to the type-flow checker.
JSON conversion does not belong in the query, and projections are always
typed. Whatever surfaces at the membrane is serialised there; a declared
record requires no separate conversion. Return a RefundDecision and the
caller receives its fields as JSON, just as the incoming request was
deserialised into its declared input type. The record declaration also allows
the type-flow checks to continue through downstream nodes.
A search may combine nested enrichments with a projection in one statement. The projection can refer to the search alias and every declared nested name:
FIND c<Company> WHERE c.companyId = input.companyId ON_EMPTY EMPTY_LIST {
departments: FIND d<Department> WHERE d.companyId = c.companyId ON_EMPTY EMPTY_LIST
}
PROJECT INTO CompanyView { name: c.name, departments: departments }
A complete query and projection
The following example brings the pieces together in one statement. The search
finds the request's account, then the account's orders and each order's line
items; the tail then reshapes those enriched rows into a declared record. The
tail can reach the request as input, each nested list by its name, and any
depth of the nesting as a dotted path — orders.items.amount reads through
every order's items:
- type: JEQLTaskNode
id: account-summary
query: |
FIND a<Account>
WHERE a.accountId = input.accountId
ON_EMPTY RAISE_EXCEPTION
ON_MANY RAISE_EXCEPTION {
orders: FIND o<Order>
WHERE o.accountId = a.accountId
ON_EMPTY EMPTY_LIST {
items: FIND li<LineItem>
WHERE li.orderId = o.orderId
ON_EMPTY EMPTY_LIST
}
}
LET highValue := orders.items [ amount >= input.minimumItemAmount ]
PROJECT INTO AccountSummary {
accountId: a.accountId,
accountName: a.accountName,
orderCount: count(orders),
highValueCount: count(highValue),
totalItemValue: sum(orders.items.amount)
}
inputType: domain.AccountReportRequest
outputType: domain.AccountSummary
The aggregates here fold per matched account, over the paths they are given:
count(orders) counts that account's orders, sum(orders.items.amount) sums
every item of every order, and the [ … ] filter keeps the items that clear
the request's threshold. Because the entire nesting is declared in the same
statement, all of these paths are checked at build time. A misspelled field is
therefore reported during compilation. The tail supports LET, WHERE,
GROUP BY, HAVING, ORDER BY, DISTINCT, SKIP and TOP. A per-row value
such as count(orders) must be computed before GROUP BY and bound with
LET, because the rows have been replaced by groups after that stage.
Alongside the aggregates there is a small set of scalar functions for the
transforms that would otherwise force a Java task node per trivial reshaping:
lower, upper, trim, len, substring, startswith, endswith for
strings; abs, round, floor, ceil for numbers; year, month, day,
hour for timestamps, with instants interpreted in UTC. Functions may be
nested, as in lower(trim(p.name)). Null inputs remain null, while a boolean
function such as startswith(p.name, 'A') can serve as a predicate.
They belong to the projection side of a query — values, LET, the tail's
WHERE, HAVING, GROUP BY, ORDER BY, [ … ] filters — and not to the
seek WHERE that selects stored records, which must stay eligible for an
index; a normalised comparison goes in the tail.
Unprojected results
Without a projection tail a nested search has no answer to give: the frames it
builds — a list of alias-keyed maps, {a: <Account>, orders: […]} — are its
internal structure, and the engine refuses the query at build until you declare
the record you want and project into it. A nested name may ride along as a
field (List<Order> for a flat child), so nothing is lost; what changes is that
the shape leaving the node is one you named. The FOR input<T> { … } binder
ends the same way and answers one record — the shape to reach for when the
next task wants the request beside what was looked up.
A projection does not read stored records. It only reshapes the value it receives, which accounts for the synchronisation exemption described below.
Indexes
A record is findable whether or not an index covers it. Indexes change how
long a search takes, never what it returns, and an edition without
multidimensional indexing runs a gist as a scan that produces the same
answer more slowly. The practical concern is therefore performance rather than
correctness.
Declare a covering index on every search that will run against real
volume. Indexing starts once a type's committed records pass a threshold
(lounge.indexing.threshold, default 100), so on a small space the
declaration costs nothing, and on a large one it is the difference between a
probe and a scan of everything.
Index per query
The instinct from SQL is to index the table and let every query find what it
can. Here the declaration sits on the query, beside the WHERE it serves, and
the engine builds that index for that query alone. Two queries over the same
type each carry their own; a query nobody wrote has no index, and a query
nobody runs costs nothing to serve. So an index is shaped by the predicate it
answers rather than by a guess about the table, and "covering" means one
thing: every column the query's conditions touch appears in its declaration.
The choice is mechanical. If every condition in the WHERE is an equality,
declare a hash over those columns. If any condition is a range — <, >,
BETWEEN, or several dimensions at once — declare a gist over every column
the query touches: the gist folds all the equality columns into one hash key
and holds it as a single dimension beside the range columns, so one index
serves the whole predicate.
# one equality → hash on it
FIND u<User> WHERE u.email = input.email INDEXES { byEmail: hash(email) } ON_EMPTY RAISE_EXCEPTION ON_MANY RAISE_EXCEPTION
# several equalities → one hash over all of them
FIND o<Order> WHERE o.customerId = input.customerId AND o.status = input.status INDEXES { byCustomerStatus: hash(customerId, status) } ON_EMPTY EMPTY_LIST
# an equality and a range → gist over both; customerId folds into the hash dimension
FIND o<Order> WHERE o.customerId = input.customerId AND o.amount > input.floor INDEXES { byCustomerAmount: gist(customerId, amount) } ON_EMPTY EMPTY_LIST
# a range alone → gist on it
FIND r<Reading> WHERE r.value BETWEEN input.low AND input.high INDEXES { byValue: gist(value) } ON_EMPTY EMPTY_LIST
# an equality and two spatial dimensions → one gist, three dimensions
FIND s<Site> WHERE s.region = input.region AND s.lat BETWEEN input.lat0 AND input.lat1 AND s.lon BETWEEN input.lon0 AND input.lon1 INDEXES { byRegionGeo: gist(region, lat, lon) } ON_EMPTY EMPTY_LIST
A nested search declares its own the same way, on its own line, because it is its own query:
FOR input<Customer> {
recent: FIND o<Order> WHERE o.customerId = input.id AND o.placedAt > input.since INDEXES { byCustomerDate: gist(customerId, placedAt) } ON_EMPTY EMPTY_LIST
}
PROJECT INTO CustomerRecent { id: input.id, recent: recent }
Declare nothing and the engine reads the WHERE clause and chooses for you.
When in doubt between the two, declare both: a hash whose columns a gist
already covers is detected at build, never built, and named in the log. The
query reference has the folding rule in full.
Query synchronisation
The parallelism page made this point from the scheduling side; from the query side it is worth repeating, because it explains two rules that otherwise look arbitrary.
A query is the one place in a flow that depends on what other work has finished. So the engine makes sure it has:
- An event's first query waits for the previous event in its group to finish. Under overlapped admission, events process concurrently right up until one of them asks a question — and asking is where it stops and lets its predecessor complete. A query never sees half of an earlier event's records.
- A searching query may not sit inside a parallel fan-out. There the other work is a set of peers with no order between them, so there is nothing to wait for and no correct moment to run. Put the query after the fan-out. The linter enforces this; projections, which search nothing, are exempt.
Both rules enforce the same condition: a record-reading query runs only after the writes on which its answer may depend have settled.
Query placement
The first rule has a direct design consequence: an event overlaps with its predecessor only up to its first query.
So the same two nodes, in the two possible orders, are not equivalent:
work (1ms) ──▶ query the work overlaps; by the time the question
is asked the predecessor has finished anyway
query ──▶ work (1ms) the event blocks immediately; the work that
could have overlapped happens afterwards
The practical rule is: do your work first, ask your questions late. Validate, transform, call out to the world, produce your records — then query. A flow written that way overlaps freely; a flow that opens with a lookup has serialised itself before it has done anything.
This is not a licensing matter. A query-first flow is slow in the same way on every edition; reordering it is a spec change, not an upgrade.
Continuous queries
A standing query asks its question of the past and of everything that
arrives afterwards: tree.subscribe(...) turns a query you already wrote
into the definition of "interesting", and each newly committed record is
tested against it as it commits. The query reference has the subscription
API and the no-gap recipe.
Record lifetime
Memory is finite, and a long-running application accumulates. The retention rule from the top of this page is applied over time — external arrivals, exceptions, query-read types and explicit keeps stay; other values are dropped at commit — and whatever retention drops, sessions always survive: the history of what happened is not governed by record retention.
Past that, quiet groups are reclaimed by an idle timeout, and a group that never goes quiet can age its records out of a hot window into a cold SQL table: the two-lane pattern, hot in memory and cold in a plain JDBC table, each record carrying a natural id so that archiving is idempotent. The Recipes page walks the archive through end to end.
Surviving a crash is a different question from staying in memory, and Durability, Clustering and HA answers it: a write-behind log that queries never read, and recovery that restores state without re-running work.
Some rules to remember
- You never insert. Records are what your task nodes returned.
- The default retains external arrivals, exceptions, and query-read types; a chain of same-type refinements keeps only the final version.
FINDsearches,FORbinds. Sought record on the left of a condition, value you hold on the right.- A query returns a list, and empty is the empty list unless the query
ends
ON_EMPTY RAISE_EXCEPTIONor (nested)DELETE_PARENT;ON_MANY RAISE_EXCEPTIONasserts a single-record answer. - A query hands on what its text declares. A bare
FINDhands on the found records; a composed query — nestedFIND, orFOR input<T> { … }— hands on the record it projects into, and nothing else leaves the node. - Projections end in a declared record, and the membrane serialises it.
- Declare a covering index on every real search —
hashwhen every condition is an equality,gistotherwise. - A query is a synchronisation point — do your work first, ask your questions late, and never inside a fan-out.
Where to go next
- Event Groups, Ordering and Parallelism — the ordering rules the synchronisation section leans on.
- Durability and HA — surviving a restart, and what that costs.
- Deep reference:
tools/lounge/references/query.md(the full DSL, indexes, hooks, standing queries),tools/lounge/references/projection.md(reshaping in detail), andtools/lounge/references/persistence.md(retention, the advanced storage knobs, the archive, the sweep).