approach
Progressive System Design Interview
A practical approach to System Design interviews that tests solution derivation, engineering judgment, trade-offs, and realistic delivery instead of memorized hyperscale patterns.
System Design interviews are supposed to test engineering judgment.
Too often, they test something else: whether a candidate has seen the same architecture diagram as the interviewer.
“Design YouTube.”
“Design Uber.”
“Design Twitter for one billion users.”
A few minutes later, the whiteboard contains Kafka, Kubernetes, object storage, several databases, CDN, queues, caches, multiple regions, analytics pipelines, and enough infrastructure to keep several platform teams busy.
The architecture may even be technically correct.
The problem is that neither the candidate nor the interviewer has necessarily built anything remotely close to YouTube. The company itself may operate at a few thousand requests per second, while the candidate’s experience may be limited to systems serving hundreds of thousands or a few million users.
Yet the hiring decision is made around a fictional hyperscale system outside both the candidate’s real experience and the company’s actual problem space.
At that point, System Design becomes dangerously close to guessing what the interviewer expects.

I prefer a different approach.
Start with something almost absurdly simple.
Then make it grow.
Who is this interview for?
I do not consider System Design particularly useful for Junior or regular Middle engineers.
At those levels, I care more about implementation, debugging, reasoning about code, understanding the immediate technical environment, and working effectively inside an existing architecture.
For the format described here, I normally expect the candidate to be interviewing for at least a Technical Lead-level role, or to be considered for promotion into one.
The interviewer should also be at least a strong Technical Lead or someone operating at a broader technical level.
System Design has very few objectively correct answers. An interviewer who only knows one stack deeply can easily turn the exercise into:
“Would you make the same technology choices I would?”
That is not what I want to measure.
I want to see whether the candidate can independently generate ideas, make decisions under ambiguity, explain trade-offs, and take responsibility for the consequences.
I describe what I mean by technical leadership separately in Technical Leadership.
Start with the smallest reasonable system
Suppose the task is:
Build a very simple video hosting service.
Do not immediately add one billion users, global distribution, real-time analytics, petabytes of storage and five nines of availability.
Start with:
We receive 10 video uploads per day. We have around 20 video views per day. Videos can initially be served exactly as uploaded.
At this point, FTP and a basic HTTP server may be perfectly valid.
That is intentional.
I want to know whether the candidate can accept a simple solution when the problem is simple.
If the first diagram already contains Kubernetes, Kafka, Redis Cluster, Cassandra and six microservices, I have learned something important.
Not about the candidate’s knowledge of those technologies.
About their ability to control complexity.
Complexity should be earned
Once the simple system works, start changing the requirements.
Video processing now takes 30 minutes.
The synchronous request no longer makes sense. Async processing becomes relevant.
Workers sometimes fail halfway through.
Now we need durable work state, retries, idempotency, or something equivalent.
Videos can be 20 GB and users upload them over unreliable connections.
Now multipart or resumable uploads matter.
Views increase 1,000x while uploads grow only slightly.
The read path becomes a different problem from the write path. CDN and caching become meaningful.
Users are spread across several continents, and a regional outage must not make uploaded videos unavailable.
Now geography, replication and failover matter.
Product wants analytics, and later advertising requires near-real-time analytics.
Now event pipelines and potentially streaming infrastructure appear naturally.
The final architecture may look very similar to a classic “Design YouTube” answer.
The difference is that every component has a reason to exist.

Solution derivation, not solution recall
A candidate can memorize that a large video platform probably needs a CDN.
That tells me very little.
I want to know whether they can recognize when and why it becomes useful.
The same applies to Kafka.
Knowing partitions, consumer groups and replay is useful. But System Design should not primarily test:
“Do you know Kafka?”
It should test:
“What changed in the requirements that now calls for Kafka-like properties?”
There is a major difference between:
We need Kafka because we have events.
and:
We now have multiple independent consumers, need durable asynchronous delivery and replay, and the volume and retention requirements make a distributed log attractive.
The second answer demonstrates reasoning.
The first demonstrates vocabulary.
My preferred sequence is:
Requirement
→ required property
→ capability
→ technology class
→ concrete product
Not:
Kafka
→ now let us find a problem for it
Technologies should come late
Concrete products should appear toward the end of the thought process.
A weak storage discussion starts with:
I would use PostgreSQL.
A stronger one starts with the properties:
We have structured entities with relationships, transactions, secondary indexes, some document-like metadata, and moderate throughput.
Therefore:
A relational datastore with reasonable document support is enough.
Only then:
PostgreSQL is a good fit. It covers both needs, the team knows how to operate it, and I do not see enough value in introducing MongoDB yet.
And if the workload later changes significantly, that decision can be revisited.
The point is not to name the “correct” database.
It is to show why the current one is sufficient and when it would stop being sufficient.
“Why not?” is often more important than “Why?”
At higher levels, some of my favorite questions are:
Why not Kafka? Why not Kubernetes? Why not microservices? Why not now?
A candidate may correctly recognize that Kubernetes could eventually be useful.
But if the current system is two stateless containers running perfectly well on a managed cloud runtime, Kubernetes may add operational complexity without solving a current problem.
A strong answer could be:
We may need Kubernetes later if the number and variety of workloads increase or the current platform stops covering our requirements. Right now its operational cost is higher than its benefit.
That is not lack of Kubernetes experience.
It is engineering judgment.
A mature technical leader must be able not only to introduce technology, but also to protect the system from premature complexity.
Sometimes the required technology does not exist yet
There is another problem with asking candidates to reproduce the architecture of companies such as Google, Amazon, or Meta.
At their scale, those companies repeatedly encountered problems for which an adequate off-the-shelf solution did not exist.
Their internal engineering work sometimes produced technologies that later became open source projects, commercial products, or major influences on the industry.
Google operated Borg for years before Kubernetes existed. Kubernetes was not simply Borg released as open source, but many of its ideas came directly from Google’s experience running Borg at massive scale.
Amazon developed Dynamo, an internal highly available distributed key-value store. The Dynamo paper later influenced an entire generation of distributed databases, including Cassandra and eventually Amazon’s own DynamoDB service.
Facebook originally developed Cassandra for large-scale inbox search and storage before it moved into the open-source ecosystem.
Google also produced MapReduce, Bigtable, and later Spanner from internal problems that existing tools could not adequately solve.
LinkedIn created Kafka because existing messaging systems did not fit its requirements for high-volume activity streams and log-style processing.
Today these technologies, or the ideas behind them, are standard items in a System Design toolbox.
This creates a strange situation in a typical “Design YouTube” interview.
The interviewer may know ten technologies. The candidate may know a different ten.
Neither has built YouTube, and neither knows Google’s complete internal technology stack.
Yet one of them is expected to judge whether the other produced the “right” architecture.
At that point the design often reflects the candidate’s toolbox more than the problem.
If the candidate knows Java and PostgreSQL well, they may build surprisingly far with Java and PostgreSQL.
If they know Kafka, Kubernetes and Cassandra, those technologies are more likely to appear.
Progressive System Design exposes this much better.
Start simple and keep increasing the requirements.
At some point the candidate reaches the boundary of their practical knowledge.
Perhaps traffic grows from 1,000 to 100,000 requests per second and Java with PostgreSQL still satisfies the requirements.
That may be completely correct.
Perhaps a later constraint finally forces a fundamental architectural change.
The useful signal is not whether the candidate eventually recreates YouTube.
It is where their current solution stops being sufficient, whether they recognize that boundary, and how they reason beyond it.
That gives us a much more useful question:
Are the real problems of our company inside the candidate’s effective problem-solving range, or beyond it?
A fictional YouTube-scale problem will usually sit beyond both the candidate’s real experience and the actual needs of the company interviewing them.
Unless the candidate really worked at that scale, or the company is actually hiring for YouTube-like infrastructure, the exercise mostly evaluates speculation.
A progressive interview lets the problem grow until we find the candidate’s real boundary.
Then we can ask the question that actually matters:
Is that boundary far enough for the problems we need this person to solve?
Architecture is not only about scale
System Design interviews often use “scale” as a synonym for more traffic, users and data.
Real systems become difficult for many other reasons:
- stronger consistency;
- compliance;
- data residency;
- unreliable external systems;
- security boundaries;
- limited budget;
- a tiny operations team;
- legacy dependencies;
- fixed deadlines;
- migrations or acquisitions;
- limited expertise.
Any of these can change the architecture more dramatically than another 10x in traffic.
Give numbers when numbers matter
Back-of-the-envelope estimation is useful when it affects a decision.
It becomes theatre when the candidate invents arbitrary numbers and never uses them.
I prefer giving enough concrete information to reason from:
100,000 daily active users. 5,000 uploads per day. 2 million views per day.
Then later:
Views increase 100x, uploads only 3x.
Now the interesting question is not:
Can you divide 200 million requests by 86,400?
It is:
What changed in the workload, which part of the architecture is affected, and what should we change first?
Numbers should support architecture, not decorate it.
The system should evolve, not be replaced every round
A useful interview creates an architectural history.
At every stage I want the candidate to explain:
- What changed?
- Which assumption is no longer valid?
- What breaks first?
- What is the smallest reasonable change?
- What new cost or complexity does it introduce?
The process looks roughly like this:
Simple problem
↓
Minimal design
↓
New constraint
↓
Identify what breaks
↓
Minimal justified change
↓
Repeat
A particularly strong signal is:
Nothing needs to change yet.
Not every new requirement deserves another service.
Changing your previous decision is not failure
Sometimes I deliberately introduce a requirement that invalidates an earlier architectural choice.
The candidate may design a perfectly reasonable single-region system.
Later:
We now require an RTO below ten minutes after losing an entire region.
A good response may simply be:
That changes the architecture significantly. My previous design no longer satisfies the requirement.
Architecture is based on assumptions.
When assumptions change, architecture can change.
I am more concerned about candidates who defend an obsolete decision because it was theirs than candidates who calmly replace it when new information appears.
Then make the architecture meet reality
Technical architecture is only part of the exercise.
At some point I introduce implementation context:
You have five Java backend engineers, three React engineers and one DevOps engineer shared with another project. Nobody has meaningful production Go experience. The first production release is expected in four months.
Now I want to see whether the architecture changes.
Go may be technically attractive.
Kafka may have useful properties.
Kubernetes may be the ideal target platform.
None of that means this organization should adopt them now.

This is where trade-offs stop being theoretical.
It is easy to say:
Every technology has trade-offs.
It is harder to say:
Kafka is technically attractive here, but given our workload, operations capacity and team expertise, I would use RabbitMQ or a managed queue for now. These are the conditions under which I would reconsider that decision.
That is the reasoning I want to observe.
Estimation is part of System Design
I also like to ask for a rough delivery estimate.
Not because I expect the exact release date.
I want to understand how the candidate estimates.
A production system contains much more than application code:
- discovery and requirements;
- implementation;
- infrastructure;
- security;
- CI/CD;
- testing;
- observability;
- migrations;
- backup and restore;
- load testing;
- documentation;
- compliance;
- support and incident readiness.
The first useful question is often:
What team do I have?
A candidate may decide that lack of Go experience makes Go a delivery risk, or that limited DevOps capacity makes several new infrastructure products a bad idea.
The exact number matters less than whether the candidate understands what work exists.
If someone designs a sophisticated production system and then says:
This is “two or three days” of work.
that tells me more than another twenty minutes of architecture discussion.
Delivery decomposition matters too
For a sufficiently senior role, I also want to know whether the candidate can turn architecture into work.
They should be able to reason about:
- critical path;
- parallel workstreams;
- dependencies;
- temporary interfaces or stubs;
- scarce specialists;
- what is required for the first release;
- what can wait;
- what needs a technical spike;
- what must remain reversible.

I do not expect a perfect project plan.
I expect something plausible: what should be built first, which workstreams can run in parallel, where the bottlenecks are, what blocks the first release, and what can safely wait.
A system that can theoretically be built is not necessarily a system this organization can practically deliver.
An ownership-level deep dive
For roles that include broader technical ownership of the project, I usually continue a little further.
The core System Design interview is still the same. I simply start adding questions such as:
What team would you need?
Where are the scarce skills?
What can be delivered in parallel?
What is the critical path?
What would you postpone?
What would you buy instead of building?
What would you deliberately keep simpler for now?
What is your rough delivery estimate?
The exact estimate is not the important part.
I want to understand how the candidate arrives at it.
A person may design a technically excellent solution and still underestimate the amount of work needed to build, operate and support it.
Likewise, a technically attractive architecture may stop being attractive once we learn that the team does not have the required expertise, operations capacity is limited, or the deadline is fixed.
At this point I am no longer evaluating only whether the architecture is valid.
I am trying to understand what scope of technical responsibility the candidate is ready to own.
A strong technical lead may be perfectly capable of designing a sound solution.
Someone ready for broader project ownership should also be able to turn that solution into a realistic delivery approach under the actual constraints of the organization.
I describe these technical leadership scopes separately in Technical Leadership.
Local decisions can have a larger blast radius
I also pay attention to whether the candidate recognizes when a supposedly local decision stops being local.
Choosing DDD or internal code structure may stay inside one application.
Introducing MongoDB, Kafka, or a new cloud provider may affect infrastructure, operations, hiring, security and support.
At that point it is no longer purely an application decision.

The candidate does not need to know the organization’s exact approval process.
I simply want them to recognize the larger blast radius.
Real experience matters
I do not expect every candidate to have solved every problem in the interview.
At some point everyone leaves direct experience and starts reasoning from fundamentals.
That is useful too.
But a good interview should spend meaningful time in the overlap between:
- the candidate’s real engineering experience;
- the company’s real technical problems;
- the problem being evaluated.
A progressive interview naturally moves through different levels of complexity and helps reveal the boundary between:
“I have dealt with this before.”
“I understand the theory.”
“I am now guessing.”
That boundary is useful information.
A fictional hyperscale problem can put both interviewer and candidate into the third category from the first minute.
What I evaluate
The final architecture is not the primary output.
I care much more about the path that produced it.
Problem decomposition
Can the candidate turn ambiguity into concrete engineering problems?
Complexity discipline
Do they add complexity because a requirement demands it, or because the technology is familiar?
Technical breadth
Can they reason across application, data, infrastructure, networking, messaging and operations where necessary?
Trade-offs
Can they explain both the benefit and the cost of a decision?
Evolution
Can they recognize when an assumption is no longer valid and adapt?
Context awareness
Do team capability, budget, operations and deadlines affect the design?
Delivery thinking
Can they turn architecture into realistic implementation work?
Ownership
Can they distinguish between a technically valid solution and one this organization can realistically build, operate and evolve?
Communication
Can another engineer understand why the system looks the way it does?
There is no perfect final diagram
The interview should not have a secret architecture that the candidate is expected to reproduce.
Two candidates can arrive at different systems and both perform very well.
One may choose RabbitMQ where another chooses a managed queue.
One may use PostgreSQL JSONB where another separates document storage.
One may keep a modular monolith much longer than another.
What matters is whether the decisions follow from requirements and whether the candidate understands their consequences.
I would rather see someone defend a simple architecture extremely well than accidentally reproduce my preferred architecture without understanding why it exists.
The core principle
System Design should not test whether someone can draw Google’s architecture from memory.
It should test whether we can give this person a real technical problem and trust them to:
- start simple;
- identify what actually changed;
- add complexity only when justified;
- explain the consequences;
- adapt when assumptions change;
- account for the team and operating environment;
- turn architecture into a realistic delivery approach;
- estimate the real work;
- recognize the blast radius of their decisions;
- take responsibility for the resulting technical direction.
The final system may eventually contain Kafka, Kubernetes, multiple databases, regional failover and complex analytics.
That is fine.
But every component should have a story.
The candidate should be able to answer:
Why does this exist?
And sometimes the strongest answer is:
It does not need to exist yet.
Do not test solution recall. Test solution derivation.