Design a URL shortener: on AWS and at scaleLESSON 9.14c · 14 OF 22 IN CHAPTER
PART C / System design under constraints
Step 147 of 255
LESSON 9.14c · 14 OF 22 IN CHAPTERGUIDED READING

Design a URL shortener: on AWS and at scale

Workload assumptions and capacity decisions

These are constructed exercise assumptions. The stated workload is a design target. The local demonstration does not establish that throughput. Use the estimation constants to check units before choosing capacity.

Input or objective Calculation / consequence
30 million links/day 30,000,000 / 86,400 ≈ 347 creates/s average. Provision for the actual burst, not just this mean.
3 billion redirects/day About 34,722 redirects/s average. Model the existing 80,000/s hot-campaign case separately.
100 ms redirect p99. 30-day default expiry Keep analytics off the critical path and enforce expiration on each lookup, including cache hits.

Map the local implementation to AWS

Deployment status: local only. Running the supplied command creates no AWS resources and configures no cloud connections. The reference architecture under Design it is a proposed deployment of the completed application. Each box needs either a deployed runtime, a provisioned service or an explicitly external dependency.

Read it by following the arrows from the entry point: application code accepts the request or event, the state owner commits it, and any worker produces the later result. The table ties those roles to code and adapter work. Multiple boxes do not imply multiple Python files already exist.

An ECS resolver behind an ALB makes its in-memory connections and cache behavior explicit at sustained traffic. API Gateway/Lambda is a reasonable smaller baseline. The database owns names. A cache only accelerates reads of already-owned mappings.

Local responsibility Cloud destination and role Implementation still required
Local HTTP listener Application Load Balancer: HTTP request routing Deploy a service behind a target group, configure health checks and bounded connection/request behavior.
Application or worker process Amazon ECS: redirect and creation service Build a container and task definition. Supply configuration, task roles and graceful shutdown behavior.
Local cache, counter or coordination state Amazon ElastiCache: mapping cache Implement a Redis/Valkey adapter and atomic operations, expiry and unavailable-cache behavior. Keep the durable authority separate.
Local dictionary, SQLite records or state model Amazon DynamoDB: unique mapping store Design partition/sort keys and write a storage adapter with conditional updates or transactions. Python state and SQL are not uploaded as a database.
Local event sequence or input stream Amazon SQS: click queue Implement producer/consumer adapters, durable acceptance, batching, and retries with a dead-letter queue.
Local counters, timestamps and diagnostic output Amazon CloudWatch: operating metrics Emit bounded metrics and logs, build the named operational view and configure retention and access.

Provision resources, then connect the application

Resource or boundary Initial configuration and reason
DynamoDB mapping table Use the short code as the unique key. Conditional create. Backup enabled. TTL is cleanup, not the expiry check.
ECS resolver + ALB One ECS service with its tasks spread across two zones behind the ALB, and a bounded connection pool. Size from measured requests per task: the drawing's eight tasks hold the commercial's peak in the board's heavy test with nothing lost.
ElastiCache Use cached target/expiry/version with bounded memory. Apply expiration and abuse policy on every response.
SQS and monitoring Queue click records separately. Observe mapping-cache misses, redirect p99, blocklist age and analytics lag.

Use one disposable AWS environment for the cloud exercise. Put the named resources in infra/template.yaml or your existing IaC tool, pass resource IDs through configuration, and scope each runtime role to its own tables, buckets and queues. The diagram is a design to implement. It is not a claim that these resources have been deployed. Record the commands you used to deploy and remove the exercise resources.

For concrete provisioning commands, configuration wiring and cleanup, use the AWS foundation guide. It includes a deployable table/queue/object-storage foundation and explains which application and service adapters you still implement.

A provisioned queue or table does not make the local program use it. Configure resource IDs in the deployed runtime, replace the local adapter, and replay the same successful and failing operation against that runtime. Record the deployed commit and observable result, then remove the disposable resources using your infrastructure tool.

How big companies keep and count clicks

Every redirect is a click. At this page's numbers that is 3 billion clicks a day. So where do they live: a database, or files in S3?

Both, and each does a different job. The redirect answers first and never waits. The click goes on a queue. A worker takes clicks in batches and writes each batch twice: as lines in cheap files that keep every click (the history), and as one +N per link on a counter (the live number). Twitter and Bitly both keep these two halves, as described below.

One click, two halves: the 302 goes back to Leo at once; the click waits on click_sqs; Click worker takes a batch of 37 clicks for launch and makes one +37 on click_counts in DynamoDB, and 37 lines go through Firehose into this hour's files in S3, which Athena reads later.

What Twitter and Bitly built

Twitter, 2021. Twitter says it processes about 400 billion events a day in real time (Twitter Engineering, 2021). One of its pipelines counts impressions and engagements on Tweets, the counts its ads revenue services read. The post's table puts that pipeline at about 4 million events a second.

  • Before: a lambda architecture, two paths for the same events. A batch path ran Scalding jobs every hour over Hadoop logs on HDFS and wrote counts into Manhattan, a database. The post's table puts the batch numbers about a day behind. A real-time path read Kafka with Heron and kept live counts in a cache called Nighthawk, 10 seconds to 10 minutes behind. When Heron fell behind, the fix was to restart it, and restarts could lose events, so live counts could come out low. A query service merged the two.
  • After: a kappa architecture, one path. Each event gets a UUID and goes from Kafka to Google Pub/Sub. Dataflow drops duplicates by that UUID, within a time window, and adds the events up as they arrive. The counts go to Bigtable, about 10 seconds behind. No events are lost on a restart, and late events are counted.
  • They kept the raw events anyway. Twitter also wrote raw events into BigQuery to measure duplicates, and loaded the old batch counts there to compare every key. For the Tweet interaction stream, over 95% matched exactly. Most of the rest were late events the old batch had dropped.
Twitter before and after 2021: two paths, batch through Scalding into Manhattan about a day behind and real time through Heron into Nighthawk 10 seconds to 10 minutes behind, merged by a query service; then one path, Kafka, an event processor that adds a UUID, Pub/Sub, Dataflow that drops duplicates and adds up, counts in Bigtable about 10 seconds behind, raw events in BigQuery.

Even after removing its batch path, Twitter still wrote raw events somewhere it could query them, and that is how it checked the live numbers.

Bitly. Bitly's engineers wrote in 2012 about writing metrics inside the request: if the metrics system goes down, do your requests hang or fail? Their answer was to put each event on a local queue, NSQ, and let another process write it on (Bitly Engineering, 2012). NSQ copies every message to each "channel" of a topic, and NSQ's design document draws a clicks topic copied to three channels: metrics, spam_analysis and archive (NSQ design).

The archive is kept on purpose. In Bitly's words, it is "kinda hard ahead of time to think of all the types of metrics you might want to collect", and the archive is a "plan z" for when data downstream is lost. A 2014 write-up of a Bitly talk adds that each redirect sends one message to an archive service that saves it to HDFS and S3, to real-time analytics, to longer-term history analytics and to an annotation service, at 6 billion clicks a month (High Scalability, 2014). That last part is a third party's notes of the talk, not Bitly's own text.

One click in S3, and the same click in DynamoDB

In S3, a click is one line in a file. The line holds an event id, the code, the time, the referrer, the country and the device: about 150 bytes. Firehose gathers lines and writes a file when 5 MB or 300 seconds is reached, whichever comes first. Those are its defaults, and you can change both (buffering). It puts each file under a folder for the hour, in UTC, such as 2026/10/04/13/ (object names). Turn on its newline delimiter, or end each record with a newline, or the lines run together (delivery). A line is written once and never changed.

In DynamoDB, the live number is one item per link in click_counts: the key, a count, and when it was last updated. Click worker adds up its batch first. If the batch holds 37 clicks for launch, it makes one update, count + 37, not 37 updates.

The same click kept two ways: in S3 one line in a file under clicks/2026/10/04/13/ with event_id, code, ts, referrer, country and device; in DynamoDB one item in click_counts with code, count and updated_at, changed by one write of count + 37 for the batch.

Wins and losses, at this page's 3 billion clicks a day. Prices are for us-east-1, read on 2026-10-04 from AWS's pricing pages for DynamoDB, Firehose and Athena, and from AWS's price list file for S3, which carries the rates the S3 pricing page shows once you pick a region.

The history: lines in S3 The live number: click_counts in DynamoDB
What a write costs Firehose charges $0.029 per GB, but rounds each record up to the next 5 KB. One click per record: 3 billion × 5 KB ≈ 14,300 GB, about $415 a day. 33 lines per record: about 430 GB, about $12.60. S3 then charges $0.005 per 1,000 files written; a 5 MB file holds about 35,000 lines. $0.625 per million writes of up to 1 KB. One +37 is one write. One write per click would be 3 billion × $0.625 ÷ 1 million = $1,875 a day.
How fresh the number is Minutes: a click lands when Firehose writes its file, up to 300 seconds by default. Seconds: as long as the worker's batch takes.
Questions you can ask later Any question the lines can answer: by country, referrer, device, hour, bot or not. Athena charges $5 per TB it reads. A whole day of lines is about 450 GB, about $2 to read. Only "how many so far". Nothing else was kept.
A viral link (the hot key) Nothing to fight over: its lines spread across many files. Every update goes to one item. One partition takes at most 1,000 writes a second (DynamoDB docs). This page's hot campaign, 80,000 clicks a second, written one by one, is 80 times that. In batches of 1,000 clicks it is 80 writes a second.
Keeping a year 450 GB × 365 ≈ 165 TB. At about $0.022 per GB a month that is roughly $3,600 a month, before compression. The counters are one small item per link. Keeping every click here instead costs $0.25 per GB a month, about 11 times S3.
Fixing a counting bug Replay: run the count again over the lines with the fix, dropping repeated event ids. You cannot. A counter only holds the total. AWS says a counter update is not idempotent: a retried update adds again (atomic counters). And if Click worker fails mid-batch, Lambda puts the whole batch back on the queue by default (Lambda with SQS), so a +37 already made can be made again.

What the raw clicks are for

Each of these needs the clicks themselves. A total cannot do any of them.

  • Charts for the link's creator: referrers, countries, devices and time of day are a GROUP BY over the lines.
  • Bot and fraud filtering: a bot shows up as a pattern across many clicks. You find it in the history, then count without it.
  • Abuse and spam takedowns: a sudden burst from one referrer or one country is visible in the lines. NSQ's design document draws a spam_analysis channel beside the archive on its clicks topic.
  • Billing by usage: if customers pay per click, they can ask to see the clicks behind the bill.
  • Recounting after a bug: run the count again over the lines. Twitter checked its new live counts against counts made from its stored events.
  • Questions nobody has asked yet: Bitly's reason for the archive.

The answer to give

Keep both, fed from one queue. The redirect never waits. The history is files nobody edits: Firehose into S3, read with Athena. The live number is a batched counter: one +N per link per batch, in DynamoDB. When the two disagree, the files win, because you can count them again.

In an interview:

  • Asked for analytics: "Each click goes on a queue and the redirect does not wait. A worker writes every click as a line into S3 through Firehose, in hourly folders, and we query it with Athena. Those files are the record, so we can always recount."
  • Asked for a live count: "The same worker adds up each batch per link and makes one atomic add per link in DynamoDB. It is seconds behind, never on the redirect path, and a viral link costs one write per batch, not one per click."
  • Asked for both: "Both, from the same queue. The counter is fast but can drift on retries. The files are the truth, and the counter can be checked against them, the way Twitter checked its live counts."

Your design is this design. The design you drew (SQS, Lambda, Kinesis or Firehose, S3, Athena) is the history half, and it is right, with one change. Between Click worker and S3, use Firehose. Firehose writes into S3 by itself, with no code of yours (Firehose). A Kinesis data stream only holds the clicks until a consumer reads them: your code, or a Firehose reading the stream (consumers). On its own, nothing reaches S3. The SQS queue in the AWS table above works the same way: something has to read it. Then add the live half: the same Click worker's one +N per link in click_counts.

Not covered here: unique visitors, which need a visitor id on each line and a different kind of counter, and counting each click exactly once.

Extend the design after the baseline works

Worked follow-up: Allocate custom aliases across two regions

Two customers in different regions both request /launch. Checking each regional cache can return available twice. Replication after both writes cannot retroactively give both customers the promised alias.

Starting design Changed requirement
One database decides whether an alias is available. Both regions accept traffic, but one authority owns alias creation.

Revised architecture. Follow the changed responsibility and failure path below. This is a design to implement. The supplied local example does not provision these components.

Diagram: Worked follow-up: Allocate custom aliases across two regions

What to implement. Route all alias mutations to a designated home-region allocator. Store alias, tenant, destination and request identity in one conditional create. Regional resolvers may read cached mappings under the existing expiry and takedown rules. During loss of the allocation authority, keep safe redirects working and return a retryable failure for new aliases. Promotion requires fencing the old allocator and proving the promoted state contains acknowledged reservations. A second DynamoDB table or DNS change alone supplies neither guarantee.

Walk through the result. Submit launch from regions A and B simultaneously. One tenant receives success and the other a conflict. Next disconnect B from the allocator. B must not allocate launch locally. Record the added cross-region allocation latency and the availability sacrificed to retain uniqueness.

At multi-region scale, choose one owner for alias allocation or a globally consistent authority. Do not let two independent regional caches reserve the same alias. Quantify added write latency and what a region does during loss of the allocation authority.

Additional design cases, alternatives and original source notes

This is a commonly listed system-design interview prompt with a concrete practice contract. Assume 30 million new links/day, 3 billion redirects/day, and redirect p99 below 100 ms. Start with redirect correctness. Make analytics asynchronous. Clarify service guarantees and a first version before filling the board with services.

01 · Try this input

Input / starting state
New link
Expected result
Return a durably unique random code. Chosen aliases are guessable.

Input / condition: Create a URL with a 30-day expiry

02 · Try this input

Input / starting state
Alias collision
Expected result
One conditional create wins. The other gets a conflict.

Input / condition: Two tenants request launch

03 · Try this input

Input / starting state
Expired link
Expected result
Return a defined not-found/expired response. Never redirect stale cache.

Input / condition: Resolve after expiry

04 · Try this input

Input / starting state
Hot link
Expected result
Serve a safe cached mapping without losing expiry or abuse controls.

Input / condition: One campaign gets 80k redirects/s

Think from the contract to the boxes

A code is an identifier, not proof that the row was created. Allocate with uniqueness enforced by the durable store, then publish the mapping to caches. Redirects are reads. Clicks are events, so a slow analytics write must not block the redirect. This brief promises no new redirect after the expiry decision time. Cache (target, expires_at, version), but check now < expires_at on every resolution, even on a cache hit. TTL eviction is storage cleanup, not the check. Use 302 with Cache-Control: no-store and disable redirect-response caching in the CDN. Cache mapping data inside the service instead. An optional edge implementation must enforce the same timestamp and clock-skew policy before every redirect, not just when filling its cache.

Check with values: warm a mapping expiring at 10:00:00. Leave it in cache and resolve at 10:00:01 → 410. Repeat through the browser/CDN and confirm neither reuses an old redirect. A redirect already sent before expiry cannot recall a destination the client learned.

Abuse is a separate clock: choose a five-second maximum age for the deny-list snapshot in this exercise. A blocked code wins over a cached mapping. An older or unavailable snapshot fails closed. Test a takedown with invalidation paused. Random generated codes reduce enumeration but do not authorize a private link. Require authentication and ownership checks for private resources.

First diagram: Trace alias creation through uniqueness, cache fill, redirect, expiry invalidation, and click aggregation.

AWS service / general role Why it fits this design Alternative and when it fits better
Amazon API Gateway / request entry Authenticate link creation and shape redirects. ALB + ECS when custom redirect handling and connection reuse matter.
Amazon DynamoDB / code mapping store Conditional put gives the chosen code one owner. Aurora PostgreSQL for relational ownership and reporting queries.
Amazon ElastiCache / redirect cache Serve hot code-to-target lookups cheaply. CloudFront plus an edge resolver only when each response enforces expiry and takedown. Ordinary cached redirects weaken this contract.
Amazon SQS / click queue Move click counts off the redirect path. Kinesis when event time and replay matter more than simplicity. A stream only holds the clicks until a consumer reads them: it does not write files by itself.

Service choice follows the contract: the box label gives the generic job, while the table explains the AWS product and a reasonable substitute. Name which component owns durable truth, where retries happen, and the guarantee each managed service does not provide by itself.

Pressure-test the design

Follow-up: Two requests race for one human alias. Follow which conditional write wins and why a cache cannot arbitrate.

A cached target was changed for a phishing report. Bound invalidation delay, add a deny list that wins over cache, and explain the emergency kill switch.

Aliases become a cross-region namespace. Choose a home for uniqueness, define collision behavior during partitions, and bound how many codes can be lost or allocated twice.

Practice artifact: Trace alias creation through uniqueness, cache fill, redirect, expiry invalidation, and click aggregation. Then trace every row in the table, draw one failure, and state what the customer observes. Suggested rehearsal: 35 minutes design, 10 minutes to challenge the guarantees.

Evidence and origin: The current community interview-question catalog lists user-submitted URL-shortener reports across companies including PayPal, Microsoft, and JPMorgan Chase. Individual report dates are not shown. The entry does not show the interview date and is not a verified company rubric. The prompt contract, workload, outcomes, diagrams and solution here are original practice material. Treat company tags as reported sightings, not a prediction of your interview loop.

Interview report listing: Open the community question entry.

Sources and further reading · 18