The research method: how the URL problem's numbers, prices and "how big companies do it" were found and cited
Four kinds of research went into the URL problem. Each has a report on the box worth copying as a template.
| What | Report (template) | What it fed |
|---|---|---|
| Traffic and capacity | ~/scratch/heavy-test-research/numbers/REPORT.md (+ sanity.mjs) |
heavy.js K and PROBLEMS |
| Prices | ~/scratch/heavy-work/pricing/REPORT.md (+ prices.json, build_prices.py, examples.py, raw/) |
heavy-cost.js, tools/prices.mjs |
| Interview concepts | ~/scratch/concepts/REPORT.md |
the part settings (sharding, batching, TTL...) |
| How big companies do it | ~/scratch/guide-bigcos/DONE.md (sources section) |
the page's "How big companies keep and count clicks" |
| The interviewer | ~/scratch/heavy-work/r4/K2/research.md, docs/backlog/interview-simulator.md |
site/interviewer/ |
Traffic and capacity
- Pick the scale from the page, then anchor it to the world. The page's own numbers (3 billion redirects a day, an 80,000 a second hot campaign) set the High peak; Medium was anchored to Bitly's published rate, Low to Wikipedia's hourly trough-to-peak ratio. Everything that was a guess is listed separately under "My estimates" (the push, the share on one code, every half-life). High's climb is an estimate and says so in the source comment.
- Every capacity number gets a kind code and a numbered source: H hard quota, S soft (raisable) quota, D documented
behaviour, B vendor benchmark, R rule of thumb.
Kinheavy.jscarries the code and the source number on every line. A new number is appended to the report's source list, never slipped into the code alone. - Run a scratch model before building. The report ran a small fluid model (
sanity.mjs) of nine designs at three levels and wrote down what breaks first in each. Those nine became the unit tests' expected outcomes, both ways (test/heavy-designs.mjs). Write the table of "this design fails at this level for this reason" before the view exists. - Interview depth. What breaks must be what a bigger machine cannot fix: one hot key, one row's lock, a hard quota, the time scaling takes. Quotas that are only a ceiling are assumed raised (decided with Bruno on 2026-09-28).
- Say the caveats. The report names what AWS does not publish (no latency figures for ALB, REST APIs or CloudFront hits), which figures are secondary (Coinbase's ad, TechEmpower), and which come from a workshop rather than reference docs.
Prices
- Use the AWS Price List files, not the pricing pages. Most pricing pages load their price tables by script, so a plain
fetch shows the rules but not the numbers. The Price List files are machine-readable and versioned: name the file and
its version for every price (
AmazonS3/20260926015512/us-east-1). - Prove the two agree. The report reran eleven worked examples from AWS's own pricing pages with Price List numbers,
and all eleven came out exactly. Keep the arithmetic in a script (
examples.py) that ends with "checks failed: none". - State the conventions once: us-east-1, on-demand, first tier; free tiers noted, never applied; a 30-day month of 720 hours; GB as 2^30 bytes; the date read.
- Measure what you can. The bytes of a redirect were measured with curl against real shorteners, and the TLS
handshake with
openssl s_client. - List what is not confirmed in its own section: TLS bytes as data out (undocumented); API Gateway's own 429s unbilled (the page that said so now redirects); DynamoDB throttles unbilled (not stated).
- Keep the board honest to the research:
site/designboard/tools/prices.mjschecks every price the board uses againstprices.json, and the research's priced design at the peak must come out exactly inheavy-cost.js.
How big companies do it
- Primary sources first: the company's own engineering post or design document (Twitter Engineering, 22 Oct 2021; Bitly Engineering, 2012; NSQ's design document). Re-check every number yourself. Iris's summary in a chat is not a source.
- A third party is labelled on the page. High Scalability's 2014 write-up of a Bitly talk is cited as "a third party's notes of the talk, not Bitly's own text".
- Show a conflict rather than resolve it. Twitter's post says the batch ran hourly; its table puts the batch about a day behind. The page gives both.
- Write only what the source says. "Billing by usage" is written as a condition ("if customers pay per click") because no source said Bitly bills that way. "Not covered here" names what the section leaves out.
- Every price on the page carries where and when it was read, with the arithmetic shown at the page's own numbers.
- End with the answer to give: what to say in an interview when asked for each thing, and "Your design is this design": the reader's likely design named as right, with its one change.
Traps that cost time
- A page that hides its prices until a region is picked. The S3 pricing page renders its table only after you choose a region. The page cites the Price List file instead, and says that it carries the rates the S3 page shows once a region is picked.
- Sources that block bots. blog.x.com sits behind a Cloudflare challenge; the Twitter post was read in full from its Wayback Machine copy (2024-09-17), figures included, and the live link kept on the page. A link checker may flag it.
- The fetch tool summarises. The concepts research notes that WebFetch summarises pages. For an exact number, fetch the raw file (a Price List CSV or JSON, a page's raw HTML: the Fargate task-size table was read that way) and quote from it.
- Paywalls and partial pages. Some pattern pages show only their headings; some books were read only to their contents. Say so in the report's limits.
- Secondary figures inside primary-looking articles (press quotes of a CMO, a blog's benchmark summary): mark them secondary.
- Doc pages that moved: say "unconfirmed" when the page that said it now redirects.
- A model's numbers. Sonnet may explain a run, but a sentence with a number the run did not give is dropped. Haiku writes only what the reader said. A number on the page never comes from a model.
The interviewer
- Read the vendor's own docs first, whole. TypeSafe's docs were downloaded page by page into
~/scratch/typesafe-docs/(llms.txtis the index), and the backlog file records what Jev is before anything was designed around it. - Build the evaluation set before the answers. 150 or more questions the way people really ask (casual, compound, typos, follow-ups that lean on the previous turn), each labelled with the answer that should come back, including every question Bruno asked that missed. Measure, change, measure again, per category. The URL interviewer went from 61.4% to 99.6% on 259 questions this way.
- Spend carefully. Every real Jev call is on Bruno's key and shows on his dashboard: replay stored answers, never loop
real calls. Say it's threshold was measured on
say.eval.json(162 phrasings, 71 held out) withsite/interviewer/evaluate_say.py; measure again only when the menu's labels change.
Real model outputs as fixtures
- Haiku's writer and drawer are tested on real answers, captured once with the key from SSM
(
/soulful/iris/anthropic, never printed, never written to a file or a log), each stored with the words, the state it was asked on, the tool input as it came back and the sha256 of the exact request. Change the prompt or the tool and capture again. Read what the model wrote, not just the verdicts: round 2's live run passed 4 of 5 with a deny list written as a plain read and a stray endpoint.
The voice
- ElevenLabs bills per character. Run
site/tools/narrate.py --dry-runfirst; a run over--cap(5,000 characters) is refused before the key is read. Make three lines, check them all the way (Deepgram's words against the text, length against the step, loudness, silence at each end), then the rest. You cannot hear: Deepgram is your ear. The words helper's brief of 2026-10-04 put the period's remaining budget at about 47,000 characters.