Many freshers freeze in this round because they have never built anything that serves millions of users. The interviewer knows that. What they want to see is whether you can take a vague problem, ask the right questions, make sensible assumptions, and explain why you picked one option over another. That is a learnable skill.
Below is a plan for a 45-minute round. The running example is a URL shortener, because it is small enough to finish in time and still touches every major topic.
The 45-minute plan at a glance
- Minutes 0–5: functional and non-functional requirements.
- Minutes 5–10: scale assumptions and back-of-envelope numbers.
- Minutes 10–15: API and data model.
- Minutes 15–25: high-level design on the board.
- Minutes 25–40: deep-dive into one or two components.
- Minutes 40–45: failure handling, bottlenecks and a short summary.
Treat the timings as a guide. If the interviewer steers you somewhere, follow them, but watch the clock so you never spend 20 minutes on requirements and leave the design half-drawn.
1. Clarify functional and non-functional requirements
Functional requirements are what the system does. Non-functional requirements are how well it must do it: latency, availability, consistency, durability. Ask about both before you draw a single box.
"Let me confirm scope first. A user submits a long URL and gets a short one back, and opening the short URL redirects to the original. Do we need custom aliases, link expiry, or click analytics? Do users need accounts?"
Then state the non-functional side as assumptions the interviewer can correct:
"I'm assuming redirects must be fast, say under 100 milliseconds, and highly available, because a dead short link is broken everywhere it has been shared. It's fine if a new link takes a second or two to work everywhere, so I don't need strong consistency there."
Write the agreed list in a corner of the board. Cut scope out loud: "I'll keep analytics out for now and come back to it if we have time." That one sentence shows you can prioritise.
2. State scale assumptions and do the arithmetic
Sometimes the interviewer gives you numbers. If not, propose them and ask if they are reasonable. Round aggressively; nobody wants long division on a whiteboard. Here is the worked example for the URL shortener:
- Writes. Assume 100 million new short URLs a month. A month is about 2.6 million seconds (30 × 86,400 = 2,592,000), so 100 million ÷ 2.6 million ≈ 40 writes per second.
- Reads. Assume links are read 100 times more often than they are created. That is about 4,000 redirects per second on average. Plan for peaks of 2–3×, so roughly 10,000 per second.
- Storage over 5 years. 100 million × 12 months × 5 years = 6 billion URLs. At about 500 bytes per record (long URL, short code, timestamps, owner id), that is 6 × 10⁹ × 500 bytes = 3 × 10¹² bytes, or about 3 TB. With three copies for replication, about 9 TB.
- Short code length. Using base62 (a–z, A–Z, 0–9), 6 characters give 62⁶ ≈ 56.8 billion codes, comfortably more than 6 billion. Five characters give only about 916 million, which is too few.
- Cache size. 4,000 × 86,400 ≈ 350 million redirects a day. Caching an entry for 20% of those is 70 million × 500 bytes ≈ 35 GB. That is an upper bound, since popular links repeat, and it fits in memory on a small cache cluster.
The numbers only matter if you use them. Say what they tell you:
"So this is read-heavy, total data is modest, and 40 writes a second is easy for one database. The hard part is serving around 10,000 reads a second with low latency, so caching matters more here than write throughput."
3. Define the API and the data model
Two endpoints are enough for the core:
-
POST /urlswith the long URL and optional alias and expiry. Returns the short code. -
GET /{code}returns a redirect to the long URL, or 404 if the code does not exist or has expired.
There is a small tradeoff to mention here. A 301 (permanent) redirect lets browsers cache it, so repeat visits may never reach your servers, which cuts load but means you cannot count those clicks. A 302 (temporary) redirect sends every click through you.
For the data model, one table is enough: code as the primary key, then
long_url, created_at, expires_at and
user_id. Then say the access pattern out loud: "Almost every read is a
lookup by code, with no joins." That sentence drives your database choice later.
4. Draw the high-level design
Keep it to five to seven boxes: client, load balancer, stateless app servers, a cache such as Redis, the database, and a way to generate codes. Label the arrows, and walk one read and one write through the diagram as you draw.
"A redirect hits the load balancer and goes to any app server. They hold no state, so I can add more as traffic grows. The server checks the cache first; on a miss it reads from the database and fills the cache."
For code generation, give two options and pick one. Hashing the long URL and taking the first few characters needs collision handling. A counter encoded in base62 never collides, but needs coordination; a common fix is to hand each app server a block of counter values so it can issue codes without asking anyone. Sequential codes are also guessable, which you should mention if links might be private.
5. Deep-dive where the interviewer points
The interviewer will usually pick one or two areas. Have a clear, reasoned paragraph ready for each of these:
- Caching. Cache-aside: read the cache, fall back to the database, then populate the cache. Evict with LRU and set a TTL. Think about what happens when a viral link's entry expires and thousands of requests miss at once.
- SQL vs NoSQL. Give reasons, not fashion. "Reads are single-key lookups, there are no joins, and we expect billions of rows, so a key-value store fits and shards easily. PostgreSQL with partitioning would also handle this size; I'd choose it where I need transactions, such as user accounts and payments."
- Partitioning and sharding. Shard by a hash of the code so data spreads evenly. Range-based sharding on sequential codes would send all new writes to one shard. Consistent hashing reduces how much data moves when you add a node.
- Replication. One leader takes writes, followers serve reads. Replication lag means a brand-new link may not be on a follower yet; we already agreed a short delay is acceptable.
- Queues. If analytics is in scope, do not write a database row on every redirect. Publish a click event to a queue such as Kafka and let consumers aggregate it. The redirect stays fast; the counts can lag by a few seconds.
6. Say the tradeoffs out loud
Every choice gives something up, and the interviewer wants to hear that you know what. A simple pattern: what you chose, why, what it costs, and when you would change it.
"I'm using asynchronous replication because it keeps writes fast. The cost is that if the leader dies, we might lose the last few writes. For short links that's acceptable. For payments it wouldn't be."
You do not need a perfect answer. A reasonable choice with the tradeoff stated clearly scores better than the "right" choice with no explanation.
7. Handle failures and name the bottlenecks
Go through the diagram box by box and ask what happens when it breaks:
- An app server dies. The load balancer's health checks remove it; other servers carry the load.
- A cache node dies. Its traffic falls through to the database. Can the database handle that share of 10,000 reads a second? Read replicas give headroom.
- The database leader dies. Promote a follower. With asynchronous replication, recent writes may be lost.
- One link goes viral. A single hot key can overload one cache node. Keep a small in-memory cache on each app server for the hottest codes.
Finish with a 30-second summary of the design, the main tradeoff, and what you would build next with more time.
What a fresher is expected to show, compared with an experienced hire
The bar is different. From a fresher, interviewers generally look for structure and reasoning: you clarified requirements, your arithmetic was correct, your diagram was clean, you knew what a cache, a replica and a queue are for, and you could compare SQL and NoSQL with reasons. You are not expected to know tuning details of specific databases or to have stories from production outages.
An experienced hire is expected to drive the whole conversation without prompting, go deep on two or three components, bring up monitoring, cost and migration, and draw on failures they have actually seen. If you are a fresher, aim for breadth with clear reasons, and be honest at the edge of what you know:
"I haven't used consistent hashing in a real system, but my understanding is that it limits how many keys move when a node is added. Is that the direction you'd like me to go?"
Common mistakes
- Drawing before asking. You end up designing a system nobody asked for.
- Estimates that go nowhere. Calculating QPS and then never using it is wasted time. Tie each number to a decision.
- Naming tools without reasons. "I'll add Kafka" is not an answer. "I'll add a queue so analytics writes don't slow redirects" is.
- One big box called "backend". Break it into components you can discuss separately.
- Skipping the data model. The tables and access patterns decide most of the storage choices.
- Going silent while drawing. Narrate. The interviewer cannot grade what you are thinking.
- Treating hints as criticism. A follow-up question is usually a pointer to where the interviewer wants depth.
How to practise
- Work through eight to ten classic problems. URL shortener, pastebin, rate limiter, chat app, news feed, notification system, file storage. Do each one end to end, out loud, with a 45-minute timer.
- Drill estimates separately. Memorise a few anchors: 86,400 seconds in a day, about 2.6 million in a month, about 31.5 million in a year. Then practise QPS and storage until they take under two minutes.
- Learn one building block at a time. For caches, queues, replication and sharding, be able to say what each does, when to use it, and what it costs.
- Practise with someone asking follow-ups. A friend works if they keep asking "why?". You can also sit a system design round on mangoose.tech: it opens a whiteboard where you draw while talking to Meera, who asks follow-ups the way a panel does. The gap report afterwards notes whether you stated scale assumptions, which components you drew, only discussed or missed, which tradeoffs you covered, and whether you discussed the data model and failure handling.
Practise a system design round out loud
Draw your design on a whiteboard while Meera asks follow-ups, then read a report showing what you covered and what you missed.