Dry run: the ticket inventory problem under this playbook (sketched, not built)
The page: curriculum/03-production/01-system-design/problems/ticket-inventory.md, "Reserve concert seats with expiring
holds". Chosen because it is not a read-heavy problem: its hot path is a contested write, a hold is a lease that expires,
a retry must not create a second hold, and a group of seats is held all or none. It was also the heavy test research's
runner-up ("the purest flash crowd", ~/scratch/heavy-test-research/numbers/REPORT.md §1). Realtime chat would be a
bigger step: an answer today goes only to its caller (nothing pushes to Ben), and nothing in the runner or the heavy model
holds a connection open (API Gateway lists "websocket" among its search words, but its kind is a request gateway).
Every fact below is the page's own unless marked "to research" or "to decide".
Hour 0: the problem sheet
- Scenario: a stadium tour goes on sale; fans hold seats while they pay. On its own line: at 10:00 the sale opens, and 60,000 attempts a second compete for 30,000 seats for 20 seconds (the page: 1.2 million attempts for 30,000 seats).
- Requirements the interviewer reveals: a hold has an owner and expires after five minutes; only one buyer gets a valid hold on a seat; checkout turns a live hold into a sale for the same owner; cached availability never authorizes a sale; expiry is a checked timestamp, not when cleanup runs; a lost reply retried with the same request ID returns the original hold; up to four seats are held all or none; a payment after its hold expired is reconciled (refund or reacquire), and a sold seat never goes back on sale; extra attempts are refused clearly, with 429 and Retry-After.
- The proven answer (each to be sourced before the board is built): the seat's conditional write is the moment of ownership (a DynamoDB condition expression, or a row with a status check); a hold is a lease whose expiry is checked on every transition; the request's result is stored under its request ID; a group is one transaction; admission (a waiting room or tokens) sits before the store and grants a try, never a seat. To research for "how big companies sell seats to a crowd": the ticket sellers' own engineering writing on their queues, AWS's own documentation of a virtual waiting room, DynamoDB's transaction limits and costs. Nothing here is claimed yet.
- People: Ana and Ben (the page's names), the payment provider (an outside service), the repair worker.
The tests (ticket-inventory.tests.json), easiest first
| id | The scene, in his register | What has to happen (the check) | The page's sentence |
|---|---|---|---|
ana-holds |
Ana grabs A-17 and pays within five minutes. | Ana is told the seat is hers, then that it is sold. | "Return a hold ID and expiry time." / "Checkout converts a live hold to sold under the same owner." |
seat-race |
Ana and Ben both grab A-17 at the same moment. | One of them holds it, the other is told it is taken; never both. | "only one should receive a valid hold" |
hold-expires |
Ana holds A-17 but pays a second after her five minutes, with the cleanup paused. | Ana's checkout is refused, and Ben can hold A-17 right after. | "Expired holds still fail checkout because the write checks expires_at." |
lost-reply |
Ana's reply gets lost on the way and her phone asks again. | She gets the same hold back, and A-17 is held once. | "Lose a reservation response and retry" / "The original hold ID returns." |
four-seats |
Ana asks for A10 to A13 just as Ben takes A12. | Ana gets none of the four, Ben keeps A12, A10, A11 and A13 stay free. | "If one condition fails, reserve none" |
A sixth for later: Ana's payment lands after her hold expired and Ben now holds the seat; Ana is refunded, Ben keeps it, the seat is never sold twice.
Against today's runner: only seat-race would run, and only if the seat id were sent as the body's alias. It would
still play in the URL's words: Maya and Sam, a "poster link", "the name is taken". The other four need the runner's world
made general first (generalize.md §1, §2):
- Two keys per request. A hold is stored by seat, and its result by request ID. Today a request has exactly one
code. - A conditional write "if free or expired". Today's "if not exists" refuses whenever the seat has a record, expired or
not.
WRITE_MODESis data inflow-view.js, but each mode needs its runner semantics too. - A write over several keys, all or none. No verb does this today.
- A transition on an existing record (checkout: held to sold, same owner). Today POST and PUT always mean "create a link", DELETE means "take it down", GET means "open it, and count a click".
- New scenario events: pause the cleanup, lose a reply, a late payment. Today only
cache-drops,pause-invalidationandslow-clicksexist.
What the API can name
- Resources: Seats table (a database: seat, status, hold id, owner, expires); Requests table (a database: request id, the result; or the same table: to decide); Seat map cache (a cache: the page says cached availability may be shown but never authorizes a sale); Repair queue (a queue); Repair worker (code); Payment provider (an outside service); API service (code).
- Verbs that fit as they are: read; write if not exists; check the owner; check the expiry; check the rate limit; send; call the provider and wait, with a failed way; answer.
- New: the "if free or expired" mode; "together" (all or none); reads and writes by request id; DELETE meaning "release the hold" rather than "mark it blocked"; outcomes "sold", "released", "refunded"; statuses meaning seat things (409 "the seat is taken", 429 "wait your turn").
The board (ticket-inventory.board.json)
heading"Design it" (the page needs the new top:# Design a ticket sale, the scenario,## Design it); title "Seats under a flash crowd"; brief: "Ask the interviewer what to build, write the API, then draw it and send it the sale's traffic with Heavy test."requirements: Hold a seat (one buyer, an owner, five minutes); Buy it (a live hold, the same owner); Retry safely (the same request ID, the same hold); All or none (up to four seats together).expectations: 60,000 attempts a second for 20 seconds; never two owners of a seat; extra attempts told to wait, never left hanging; a latency goal for a hold. The page gives none: to decide with Bruno, because the heavy test's verdict needs one.checks:arrows,reachesApi,spread;namesInDb's logic ("the database decides who holds a seat") with its own words, which needs its sentences out of code (generalize.md§5); and a new check, "admission before the store", whose logic is likeclicksOffPath(something in front of the store takes the crowd).capabilities(Break it's ids fromsim.js):save"Hold a seat",read"See the seat map".heavy:"ticket-inventory".
The heavy test
- Request kinds: seat-map reads (a CDN or a cache can answer them; the hot key is the venue's map); hold attempts (conditional writes; the hot keys are the best section's seats); checkouts (writes on held seats); payment calls (an outside service); repair events.
- Levels: Low "A quiet presale"; Medium "The fan club's presale opens"; High "General sale at 10:00": the page's own
60,000 attempts a second for 20 seconds, then the crowd drains. Low's and Medium's numbers: to research, with sources, as
the URL's were. A warning from the research itself (§1): for this problem "LOW and MEDIUM have no natural story", and
"passing HIGH means rejecting most buyers, which needs 6–8 parts". So the three-level shape may not fit every problem,
and today the level ids are fixed in
board.js(L1776): a problem should be able to name its own levels. - What should break first, from numbers already in the research: one DynamoDB item takes at most 1,000 writes a second ([21] in the numbers report); a new on-demand table starts at 4,000 writes a second and takes twice its previous peak at once ([23]); Lambda's account request quota is ten times its concurrency quota (10,000 requests/s at the default 1,000 concurrent executions), and raising concurrency raises both; a transactional write costs two write units (the pricing report). A scratch model must confirm these before any card is written.
- Next steps the board would offer: a waiting room in front of the store; pre-warming the table (to research: DynamoDB's warm throughput setting); seats spread across partitions; the seat map from a cache, with the page's warning that a cached "free" never sells a seat.
- Today's model cannot run it: a hot conditional write has no column in the
Walk(generalize.md§4).
Narration
Prefix seats. A line per situation in his register ("Ana asks for A-17. It's free, so the database holds it for her.",
"Ben asks for A-17 at the same moment. It's taken."), an introduction per scene that never gives the outcome away, made
with narrate.py --dry-run first and under the cap.
What this shows
The playbook carries over unchanged: the page shape, the problem sheet, the interviewer and its evaluation set, the tests
file's format, the scenes and their order, the narration catalog, the research method, the plan of rounds and the release.
What does not carry over is code. The runner's world, the step vocabulary and the heavy test's request columns all need
making general first, and the dry run names five specific needs in the runner alone. Built today, this problem would take
days. Built after generalize.md's rounds, it would be content plus three new verbs or modes, three new events and one new
check, which fits the hours plan.