Developing with LLMs
A Lounge application is a specification the engine can run, so a model that develops on Lounge is asked for that specification: the records, the typed steps, and the decisions the brief left open. This page reports the measured effects of specification quality, record-first decomposition and the build gate. It also reports preliminary results from using a small local model to write the tree.
Specifications
In the July experiments the same model scored at or near full marks on a crisp, itemised specification, whatever preparation it had been given, and spread from 60% to 95% on a narrative brief for the same application, at two to four times the tokens. Not one failure was syntax. Every one was a misreading of the brief: a behaviour built as a Java method instead of a flow through the tree, an "audit everything" obligation landing in one flow, a lifecycle transition never written back so the next lookup read stale state.
The results indicate that interpreting the brief is the decisive step, before any YAML exists.
The refinement phase is therefore mandatory. It turns a brief into a numbered inventory of behaviours, an acceptance test for each, the tree construct that owns each, one row per flow for every "always", and a map of which record is re-registered at each transition. A person reviews it before anything is generated. The stale-lookup defect, which every experiment arm produced, is why the last item exists.
Records
A tree is wired by types. A node accepts what matches its declared input, and the linter refuses a tree whose types cannot connect. Records are therefore the part of the specification the engine reads first.
Models behave the same way. Asked to "decompose this task", the untuned 27B model produced about forty steps per task, none saying what it produced. Asked to name each step's input and the record it produces, it produced about nine, typed end to end in five of six tasks:
2. Validate Read Integrity: RawMeterRead -> ValidatedMeterRead
4. Compute Consumption Delta: ValidatedMeterRead + HistoricalContext -> ConsumptionDelta
5. Calculate Tamper Score: ConsumptionDelta + HistoricalContext -> TamperScore
Each line is a task node with its inputType and outputType already
decided. The instruction that did the work: a subtask that produces
nothing is a validation gate or a side effect, and the specification must say
which. On briefs that named their records, the untuned model delivered the
requirements at the same rate as the tuned one. In these experiments, naming
the records had the largest effect available to the author of a brief.
Open questions
The other prompt shape worth keeping asks the model, before it decomposes anything, to list the policy questions the brief leaves open, where the rules live, at what level a computation applies, what routes where, and to answer each with a stated assumption. It surfaced three times as many decisions as the bare prompt. Those are the questions the human review exists for; the cross-cutting expansion and the writeback map are where a narrative brief silently loses most of its points.
Work then splits along the tree's own seams: the records as a task of their own, then one flow per task, ready when its purpose fits one sentence with no "and" in it.
The gate
A model's tree goes through the same gate as a person's: schema, build, strict lint. A lint finding is a failing test. Passing is the floor: the evaluation oracle scores how much of the brief a tree delivers, as structural predicates over the parsed tree, and it flags the fat handler, a task described with two verbs or with "handle", "process" or "manage". That check found a quarter of our own generated training trees to be structurally valid but semantically empty. The generator was fixed before those examples were used for training.
A local model
Lounge Author 27B v3 is now available as fused MLX 4-bit weights for Apple silicon under Apache-2.0. The package includes a compact authoring prompt, a bounded inference script, usage instructions and evaluation notes. See download and local setup.
The model drafts application trees, Java records and small functions. Its
internal evaluations are small and sensitive to the supplied skill and prompt.
Historical scores used a longer skill and different output limits from the
published quick start; they do not establish general superiority to another
model. The model card records those settings and results. Review the generated
files, run lounge build, and test the actual requirements. A structural mistake
may need a new node or a different tree, rather than another wording of a repair
prompt.
Local development
After downloading the dependencies and model, the authoring and validation loop can run on one Apple-silicon machine without a cloud inference service or API key. The weights occupy about 15.1 GB; the model card recommends at least 24 GB of unified memory as a starting point, with actual headroom depending on the prompt and output length. Application integrations may still contact external services if configured to do so.
Start with the included script's 2,048-token output limit and ask for one bounded subtree at a time. The script prints text for review and never executes generated code automatically. Keep the relevant records, subtree and exact build diagnostic in a repair prompt; long repair contexts have produced runaway output in internal runs.
Where to go next
- The skill a model is given:
tools/lounge/SKILL.md, withcheat-sheet.md,thinking-in-trees.mdandspec-refinement.mdundertools/lounge/references/. - The process skills, in order: design, plan, build, verify, under
tools/lounge-plugin/skills/. - The gate as a person runs it: Building, Operating and Securing.
- The published model and its evaluation notes: Hugging Face model card.
- The fine-tune pipeline and its oracle:
tools/finetune/README.md; the round results underdocs/product/.