The mistakes that cost the most, and what prevents each now
From the commit history (git log --since=2026-09-26 -- site/designboard curriculum/.../url-shortener*), the helpers'
briefs and reports under ~/scratch/heavy-work/ and ~/scratch/guide-*/, and Bruno's corrections in them. Most costly
first.
1. The proven answer was decided last
On 4 October four helper rounds went into one requirement, where a click is kept and counted: the retouch (an unread queue), the board explaining itself (Firehose, idle parts), the advice that ends where a click is counted, and how big companies do it. Each round answered the question the last one raised. His words that day: "i wish i could be taught while doing this, that was the point." The page's "The answer to give" was written after the board's words, not before them. The takedown was the same story that evening ("what deny list. I didnt have that nor thought i needed that.").
Guard: write the proven answer first, with its sources: for each requirement, the industry's usual pattern, where the
data ends up, and what an interviewer expects to hear. The tests, the hints, the fix cards, the heavy test's next steps and
the "how big companies do it" section are all derived from that one page (plan.md, hour 0).
2. Building before the shape was agreed
- Diagram first, when he wanted the API first. Round 2 shipped the API under the diagram (2026-09-27 morning); at 10:13 he wrote "One thing I thought I made clear. I wanted the api spec to be created first snd then the design." Round 3 redesigned the model, the runner, the writer and the panels, and took the rest of the day.
- A version switcher nobody asked for. Built at 20:03 on 27 Sep from "We should start Versioning" and removed at 07:18 the next morning: "I just wanted versioning. Versioning could be in github. Not up to the user. Up to us." (commit 8e673f8). Every release had been building all eleven tagged boards.
- The tests view, rebuilt four times in one evening: test rows, then scenes (after "I feel like im doing a checklist rather than learning", 17:33), then no rows or codes at all ("I don't wanna see any of that"), then unlockable simulations in a pop-up (19:27), then bigger and slower with a locked tip (20:52).
Guard: follow the interview's order (requirements, API, design, scale), which is now settled. For anything a reader will see that is new, show one picture or a mockup and get his yes first, as round 2's mockup did. When an ask can be read two ways, ask one line before building.
3. Telling the reader after a run instead of before
His queue sat unread and he learned it only when simulation 4 failed and the heavy test's Lambda sat at 0%: "If I knew that I need to add the event endpoint I would. why didnt an error or something showed up ... I shouldnt be uncoverijng this." Following the board's advice then moved the click one hop further, twice (click_sqs, then click_stream). On his first heavy run, three rules the diagram could not show (autoscaling only inside a group box, the click count written inline, a cliff at 3:00) sank it.
Guard: every rule a test or the heavy test enforces is said before a run, where the reader works: a caution under the step in the editor, a gate that refuses with the fix and marks the part, a requirement above the canvas. Advice must end where the thing is done (a click is counted once a step writes it into a database), never "add another taker". Every idle part says why it is idle.
4. A model that said false things
The concepts research found the heavy test claiming no click lost while a topic delivered nothing to its subscribers, a
stream split between its readers like a queue, and Retries applied whatever the arrow said (~/scratch/concepts/REPORT.md,
fixed in commit 9aec5a4). Earlier, Aurora and DynamoDB could tie where an interviewer expects a difference ("dynamo db is
the obvious winner there", 2026-09-27 10:16).
Guard: before a concept or a card is built on the model, probe the model with the smallest designs that exercise each
part kind (as check-consumers.mjs did) and fix the model first: "the model is the thing to fix, not the card" (round 3's
brief). Every model change runs old and new side by side over every saved design at every level; nothing else may move
(test/fixtures/heavy-base-digests.json).
5. Checks that could only pass
- A card check passed on an empty picture because it only asked whether an
<svg>existed; a check of the numbers on the arrows counted the speed buttons instead (commit c2f5296). api-check.mjskept being run and reported as "8 of 29, by design" for days after board 3.0 replaced the panel it tests (the README correction, commit 000ed34).- The voice's faster-speed check stayed green at 2× because every line was longer than a step there; it was moved to 1×,
where it goes red (
~/scratch/guide-narration/DONE.md). - A fixture written by the same helper encodes that helper's assumptions: round 2's worst seam was invisible to every suite because the panel's checks fed it fake results.
Guard: break each new rule once, reversibly, watch its own check go red for the right reason, restore byte-identical
by sha256, watch it pass, and list the breaks. Check against the real producer (the real runner, heavyRun, real Haiku
answers), never a copy of what you think it returns. Retire a check whose subject is gone; "N of M by design" is a smell.
6. Seams between parallel helpers
The runner put a call's time in at while the panel printed only call (round 2). One helper rewrote the named specs
while another's test still expected the old names (commit ccc07a6). The editor's menu and Say it's test built the menu
differently (325ff34). Two helpers' features met wrongly (copies and two-way arrows, 4df9d00; a new answer pattern and the
nudges, 27f8e19).
Guard: write the contract first and commit it as the base; each helper owns named files and named regions of shared files; README sections go to the helper's scratch folder, not the README. After the merge run every suite on the merged tree, then use it as a reader on the live page: that is where seams show.
7. The box ran out
Helpers died twice on 27 Sep: the 600-second background ceiling after the turn ended (10:59), and the kernel's OOM killer
with four helpers, two browsers and builds on 3.8 GB (15:17) (~/scratch/heavy-work/r4/NEXT-WAKE.md).
Guard: at most four helpers; one headless browser each at most; every full build and browser run under flock on the
round's lock (or HEAVY_MEM=1500M ~/soulful/box/heavy <cmd>); watch free -m; stay in the turn while helpers run.
8. Numbers that disagreed with each other
A report said "16,440,070 lost · analytics kept up" while its own table said 16,590,946 (the first stranger's walk). The simulations' closing line promised 80,000 redirects a second on one link while the heavy test peaked at 72,000 (commit 624e2ba). A fix card prescribed an ALB reservation that board 3.4 had removed, then said "Nothing to run" (12 of 120 cards).
Guard: one number, one source: every figure a reader sees comes from the run or from K, never typed twice. A
sentence that does not need a figure has none, so it cannot go stale. Every suggested change is applied and re-run before
it is shown, and a sweep test runs every card on every saved design.
9. The board was hard to find, and buried in text
Two first-time users on two nights did not find the board, 70% of the way down a long page (commit d204c86); his notes of 29 Sep: "There is a LOT of text before the user can start the assignment."
Guard: the page opens on the problem, the scenario and ## Design it.
10. A brief that pointed at the wrong thing
The narration brief named the heavy test; his words pointed at the simulations pop-up. The helper caught it by reading his
words against the brief (~/scratch/guide-narration/DONE.md).
Guard: every brief quotes his words exactly, and they win over the brief. A helper checks the brief against them before building.