Last updated: 2026-10-02

U
Undergraduate level
APL
Applied / Methodological — Knowledge with a 5–10 year half-life — stable practice

Chapter 12: What Design Review Doesn't Catch

For new readers

This is one instalment in an ongoing, chronological diary of building PatLang, written up in numbered "Acts" as the project actually happened, warts included. You don't need to have read the earlier instalments to follow this one, but it helps to know where it picks up: the immediately preceding instalment covers the correctness-bug swarm that turned up once the self-hosted x64 backend actually shipped as the default, alongside a real debugger and match/case pattern matching finally landing for good. This instalment opens a few days later, after a brief pause, with the project building a cleaner way to run a program in isolation — and then immediately putting that cleanliness to the test by finding bugs that only existed in the one environment nobody had been testing directly.

A direct continuation of the previous instalment — though, once again, not a continuation of its subject. The commit log goes quiet for five days after that arc closes (14–19 September 2026), a pause too short to be the nine-month gap Chapter 1 made famous, but the same shape in miniature: the work was still there when the project came back to it. What it came back to do falls into two short, related arcs — a quiet day spent giving the interpreter a genuinely clean room to run programs in, and two days spent finding out what that cleanliness was actually for.the pause lets the mind reset

Act LXXI: a room with nothing already in it

The day's first new primitive, world_run(ir, opts), does something that sounds almost too simple to be worth naming: it swaps every global store the interpreter has for an empty one, runs a program, captures its output, and restores whatever was there before — on success, on a caught error, or on an outright panic. Before this existed, running one program and then another in the same process meant trusting that nothing from the first run had leaked into the second: a stale object, an event handler still registered, a signal queue with something left in it. world_run makes that trust unnecessary by construction. The same commit's own regression test, checking every host function against the self-hosted compiler's mirror of which global store each one touches, immediately found a real gap once it started checking rather than assuming: twenty-four host functions were missing from that table — present and working under the Rust-native interpreter, invisible to the self-hosted compiler's own copy of the same list.

interp_run followed the same afternoon, moving the language's own hypothesis-testing engine onto the new clean-room primitive instead of whatever ambient interpreter happened to be lying around. That move exposed a test that had been quietly wrong in a specific way: a hypothesis script that crashed outright was being counted as a passing test, because the checker only ever looked for the literal word "false" in a script's output — a crash produces no output containing that word, so a crash and a correct rejection looked identical to the harness. A second check, run across thirty-nine example programs on both the interpreted and compiled backends, immediately found two real defects neither side had noticed alone: the native backend's IR decoder didn't recognise five of the project's own bitwise operators, so any program using them failed outright on that path, and the GOAP planner's internal bookkeeping used a hash set instead of something ordered, so two equally good plans could print out in a different step order from one run to the next.

Isolates and mailboxes, the same day's third and largest addition, are where the clean room earns its name properly. An isolate is a compiled program plus its own private virtual filesystem namespace; running one step of it means running it inside world_run and then either merging its VFS changes back — if it finished cleanly — or discarding the entire turn, VFS included, if it crashed or ran out of its step budget. Two isolates share nothing except whatever they explicitly post to each other's mailbox, each message its own file, acknowledged by id, ordered by a counter private to that one mailbox. The message queue and the signal system each gained a second implementation, purely so the same code keeps working somewhere with no real filesystem or TCP socket at all — a file-backed driver where one exists, a VFS-backed and mailbox-backed one where it doesn't — because the browser is exactly that environment, and the project's own demo pages run there.no shared memory between isolates

The same day also produced three smaller side-quests worth a sentence each rather than a section: a schema/BDD hybrid library combining declared schemas with GOAP planning and inductive logic, a symbolic-math library with ODE solving and uncertainty propagation, and three syntax-sugar demos dressing PatLang's own grammar up to read like Java, C++, and Python. None of them carry a bug worth narrating; they're mentioned here only because the same day that built isolates also, almost incidentally, built three other things.

Act LXXII: bugs the browser could see and the CLI couldn't

Two days later, a new tool turned the project's own browser demo pages into something a script could drive unattended: a headless-browser checker that serves a directory with the same cross-origin headers parslow.net itself sets, opens each page in a real Chrome or Edge over the DevTools protocol, clicks every button a page exposes, and reports anything that threw, stuck, or failed a page-specific check. It is, in effect, Chapter 1's Act VI testing discipline — don't trust a feature until something has actually exercised its real deployment surface — turned into a standing tool rather than a one-off technique repeated by hand for each new demo.

It did not take long to pay for itself. A clause as ordinary as size(books) == 3 was failing silently in the browser — no error, just false, every time, and only in the browser, never when the same logic ran natively. The cause was a quiet mismatch in how two different numeric values got tagged: the project's in-browser WASM interpreter tags the result of to_num() as "float" even when the number it was given was already an integer, while a bare digit typed directly into a clause stayed tagged "int" all the way through. Equality comparison required both sides to carry the exact same type tag before it would even look at their values — correct for telling the string "5" apart from the number 5, wrong for telling two numbers apart from each other. The fix compares any two members of the numeric tower by value regardless of which specific tag each is wearing; only a genuine type mismatch — a number against a string, a list, a boolean — still short-circuits straight to false.

The second bug was more structural. A Given/Then line combining two conditions in one ordinary English sentence — inventory contains at least one cookie and someone's balance is at least ten dollars — wasn't just failing to recognise the second condition; it was corrupting the first one. The code reading a map-shaped item name out of the sentence had no boundary on where that name ended, so it read clean past the word "and" and swallowed the entire second clause into what it thought was a single, very long item name. The fix gives that capture a boundary at the first top-level "and" or "or," then goes one step further: a compound line is split on "and" so each conjunct is checked independently, but deliberately never split on "or," because two independently-required clauses would silently demand both sides of what was supposed to be a genuine either/or. That asymmetry is tested directly — splitting on "or" is confirmed not to happen, rather than simply left unasserted — precisely because it would be the easy, wrong thing to do by analogy with "and."cf. CGA classifier

Both bugs sit underneath a new feature, inferring a declared-schema skeleton directly from a Gherkin feature file. The feature itself isn't really the story here. Neither bug would have been visible to a test that only ever ran through the ordinary interpreter or the native compiler on the command line; both needed the actual browser-and-WASM path a real user would hit, exercised by something that could click the real buttons on the real page. That is the same lesson Chapter 1's Act VI already drew from a durable queue, a REPL, a hex-grid game, and a SQL console — worth stating again here because a further pair of examples, weeks of commits later, is still turning up bugs the exact same way.

Lessons from this arc, the short version

  • A primitive that resets everything by construction is worth more than a convention asking people to remember to reset things. world_run swaps every global store for an empty one and restores it afterward regardless of how the run ended, removing an entire category of "did something leak between runs" bug rather than relying on anyone remembering to check.
  • A test that only checks for one specific failure word can count a crash as a pass. A hypothesis checker looking for the literal word "false" in a script's output had no way to distinguish a crash — no output at all — from a correct rejection: the same "absence of evidence isn't evidence of success" trap Chapter 1 already named once, in new clothes.
  • Running the same programs on every backend at once finds defects neither backend's own tests had reason to look for. A thirty-nine-program cross-backend check turned up a missing bitwise-operator decoder on one side and a nondeterministic plan order on the other, in the same pass.
  • Isolation that shares nothing but explicit messages is simpler to reason about than isolation that shares memory carefully. Two isolates communicate only by posting to each other's mailbox; a crashed turn is discarded wholesale, VFS included, rather than requiring anyone to work out exactly what state is safe to keep.
  • A bug that only exists in the browser only shows up once something actually runs the browser. Both the equality bug and the clause-corruption bug were invisible to every CLI test and visible immediately to a script clicking the real page — the headless-browser checker didn't find subtler bugs than the CLI tests did, it found bugs the CLI tests structurally could not reach at all.
  • A deliberate asymmetry is worth testing as an asymmetry, not just hoping nobody notices it's missing. Splitting a compound clause on "and" is correct; doing the same thing on "or" would silently convert a genuine either/or into a demand for both. The fix tests that the "or" case specifically does not split, rather than only testing that the "and" case does.

References: What the Literature Already Knew

Neither of this arc's two browser bugs was chased down by first consulting a textbook; both were found exactly the way the rest of this page describes, by actually running the thing and watching one specific click go wrong. They're listed here against the literature afterward, because a decades-old concurrency model and a well-known testing distinction already had the shape of what this arc built and found.

  • Isolates sharing nothing but mailbox messages — an isolate is a compiled program plus its own private state, communicating with other isolates only by posting to a named mailbox, with a crashed turn discarded wholesale rather than partially recovered. This is close to a direct implementation of Hewitt, Bishop and Steiger's original actor formalism, in which independent actors exchange only messages and share no memory1, and of the specific reliability argument Armstrong later made for exactly this shape of isolation: a process that shares nothing can simply be allowed to crash and be discarded, rather than requiring anyone to reason about what state it's safe to keep2.
  • Two bugs invisible to every CLI test, found by a script that clicked the real page — both the equality bug and the clause-corruption bug existed only in the gap between what a command-line test exercises and what a browser visitor's click actually does. Fowler's distinction between a test double and the real thing it stands in for names exactly this gap: a stand-in that behaves differently from production, in a way nobody has reason to check, is where this kind of bug lives3.

  1. Hewitt, C., Bishop, P., & Steiger, R. (1973). A universal modular ACTOR formalism for artificial intelligence. Proceedings of the 3rd International Joint Conference on Artificial Intelligence (IJCAI), 235–245. ↩

  2. Armstrong, J. (2003). Making reliable distributed systems in the presence of software errors (Doctoral dissertation, KTH Royal Institute of Technology). ↩

  3. Fowler, M. (2007). Mocks aren't stubs. https://martinfowler.com/articles/mocksArentStubs.html ↩

See also

The Journey of Building PatLang (Acts I-VI), the second instalment (Acts VII-XIV), the third (Acts XV-XXIII), the fourth (Acts XXIV-XXVIII), the fifth (Acts XXIX-XXXV), the sixth (Acts XXXVI-XLV), the seventh (Acts XLVI-L), the eighth (Acts LI-LVII), the ninth (Acts LVIII-LXV), the tenth (Acts LXVI-LXVII), and the eleventh for where this page picks up from. The isolates and mailboxes this arc builds live in self_hosting/lib/isolate.patlang and self_hosting/lib/mailbox.patlang; the VFS-backed queue and signal drivers sit in self_hosting/lib/queue.patlang and self_hosting/lib/signals.patlang, with working demos at the VFS demo and the signals demo. The three syntax-sugar demos from Act LXXI's side-quests are at Language Flavours, and the BDD-to-Z inference feature underneath this arc's two browser bugs is documented at Schemas and Scenarios.