Design a URL shortener
Your company is launching a URL shortener that anyone can use: paste a long link, get a short one to share. Its first big customers are marketing teams, who print the short links on posters and put them in ads.
On Sunday one of those links runs in a TV commercial during the big game, and we expect it to hit peak numbers the moment the ad airs.
Design it
This page's AWS diagram · compare after you have drawn yours

The redirect path
Users. A visitor opens a short link such as go.example/launch, so their browser asks this service for its destination. Visitors need no sign-in because opening a link is public; creating and managing one requires an owner.
ALB. The Application Load Balancer spreads requests over the API copies and stops sending to a copy that fails its health check, so one slow or failed copy never decides whether a visitor gets through. API Gateway would hold the ad's traffic as well, but it is paid per request: priced on the board at the ad's traffic, it costs about $1.01 per million redirects against about $0.09 for the ALB, and the whole design's cost per redirect doubles.
API service (ECS ×8, Multi-AZ). The service resolves the short name and checks expiry and blocking before it redirects, because a cached mapping can be out of date. It runs on ECS rather than Lambda because redirects are a steady, heavy load with a sudden spike on top. Measured on the board, the same API on Lambda loses 72% of its redirects on an ordinary weekday afternoon, because by default all of an account's functions share 10,000 requests a second. With that quota raised as far as it goes, it still drops part of the ad's surge, because Lambda starts at most 1,000 new environments every 10 seconds, while the ECS copies are already running.
Why ×8, and why Multi-AZ? Measured on the board with the ad's traffic, 4 copies serve at most 128,000 requests a second while up to 150,700 arrive, so 2.99% of redirects are lost. 6 copies lose none but run at 78.5% of their CPU, just under the board's 80% warning line, and 8 lose none at 58.9%. So the answer runs 8, four in each of two zones, to keep that headroom. Multi-AZ is a separate reason: the heavy test loses nothing with every copy in one zone, but the design has to keep working when a zone fails, so the copies are spread across two.
VPC. ElastiCache's nodes run inside a VPC, so the API service runs in the same one to reach them, and the VPC's rules let only the API service connect to the cache. Visitors only ever reach the ALB, so the cache never needs a public address.
Link cache (ElastiCache ×2, Multi-AZ). The API looks here first for the target, expiry and version. In DynamoDB one link is one item, and one item lives on one partition, which serves at most 6,000 reads a second, while the viral link is asked for up to 80,000. Measured on the board, the same design without this cache loses 56.6% of its redirects to the ad. The mapping is cached here rather than as a CloudFront or browser redirect, so the API can still check expiry and a fresh blocklist on every visit.
Why ×2? On the board one cache node carries the ad's load with nothing lost, so the second node is there for failure, not load. ElastiCache's Multi-AZ needs a replica in a different zone from the primary and promotes it if the primary's zone fails, so the hot links keep coming from memory instead of falling back on the table.
Links table (DynamoDB), on a miss. The API reads the durable record by short code and fills the cache, because eviction must not lose the link its creator saved. DynamoDB fits lookups by key and conditional reservations; RDS would be useful if the application needed relational queries and joins.
The 302 back. For a valid link, the API sends a 302 with the long address in Location, back through the ALB to the browser, so the browser can visit that address. Use Cache-Control: no-store because a permanent or cached redirect could keep sending visitors to an expired or blocked link.
The create path
A signed-in creator follows Users → ALB → API service → Links table, where a conditional write saves the name only if its key does not exist. Because the database decides in that write, concurrent creators cannot both own launch; the other gets a conflict, and a cache lookup cannot reserve the name.
The click path
Click queue (SQS). After answering, the API puts the click on this queue, so a slow analytics consumer does not hold up the visitor. SQS fits independent work to consume and retry; Kinesis is useful when ordered streams and replay are part of the job.
Click worker (Lambda). Lambda takes clicks from SQS in batches and writes each batch twice, because the live count and the raw history answer different questions. These short batch handlers suit Lambda; an ECS worker on Fargate is useful when the work needs a long-running process or persistent connections.
Click counts (DynamoDB). The worker adds up the batch per link and makes one atomic +N update per link, so a viral link does not need a counter write on every visit. DynamoDB gives the creator a quick total by key; querying S3 files would make them wait for the history to be scanned.
Firehose. The worker also sends the raw clicks here because Firehose gathers records and writes files into S3 without a file-writing worker of our own. A Kinesis stream would still need a consumer to deliver its records to S3, so it does not replace this arrow by itself.
Click files (S3). Each raw click is kept as a line in a file because a total cannot explain referrers, countries or a counting bug. S3 keeps that history cheaply in files we can scan and recount, while storing every click in DynamoDB would pay for a database write and item per click.
A retried batch can add to a counter twice, so keep stable event ids and drop repeated ids when recounting the files. These two writes are separate, so the arrows alone do not promise that each click is counted exactly once.