System Design Interview Questions and Answers
System design interviews ask you to design a service or explain a trade-off out loud. There is rarely one correct answer: interviewers look for clarified requirements, sensible estimates, a simple first design and a clear explanation of what you would change as load grows.
Usually asked in backend and full stack interviews, from intern / fresher to senior (5+ years).
Junior (0-2 years)
- Q1Junior (0-2 years)What is the difference between horizontal and vertical scaling?
Vertical scaling means giving one machine more CPU, memory or disk, while horizontal scaling means adding more machines and spreading the load across them.
- Q2Junior (0-2 years)What is a load balancer and what are the common load balancing algorithms?
A load balancer distributes incoming requests across several servers to improve capacity and availability, using algorithms such as round robin, least connections or consistent hashing.
Mid-level (2-5 years)
- Q1Mid-level (2-5 years)When should you use a SQL database vs a NoSQL database?
Choose SQL when your data is relational and you need transactions, joins and strong consistency; choose NoSQL when you need flexible schemas, very high write throughput or horizontal scaling for a specific access pattern.
- Q2Mid-level (2-5 years)What is caching and what are the common cache invalidation strategies?
Caching stores frequently read data in fast storage such as memory to reduce latency and database load; the main strategies are cache-aside, write-through, write-back and time-based expiry (TTL).
- Q3Mid-level (2-5 years)What is the difference between REST, GraphQL and gRPC?
REST exposes resources over HTTP URLs and is simple and cache-friendly, GraphQL lets clients ask for exactly the fields they need from a single endpoint, and gRPC uses binary Protocol Buffers over HTTP/2 for fast service-to-service calls.
- Q4Mid-level (2-5 years)What is idempotency in REST APIs and why does it matter?
An operation is idempotent if repeating it any number of times has the same effect as doing it once; it matters because network retries can otherwise create duplicate orders or payments.
- Q5Mid-level (2-5 years)How do you design a rate limiter?
Choose an algorithm such as token bucket or sliding window, store a counter per client (by user ID, API key or IP) in a fast shared store like Redis, and return HTTP 429 with a Retry-After header when the limit is exceeded.
Senior (5+ years)
- Q1Senior (5+ years)How do you design a URL shortener like bit.ly?
Generate a short unique code for each long URL (for example by base62-encoding a unique ID), store the mapping in a key-value or relational database, serve redirects through a cache because reads far outnumber writes, and record analytics asynchronously.
- Q2Senior (5+ years)What is the CAP theorem and what does it mean in practice?
The CAP theorem says that during a network partition a distributed system must choose between consistency (every read sees the latest write) and availability (every request gets a response); you cannot have both while the partition lasts.
- Q3Senior (5+ years)What is the difference between database sharding and replication?
Replication copies the same data to several servers to improve read capacity and availability, while sharding splits the data across servers so each holds only a part, which increases write capacity and total storage.
- Q4Senior (5+ years)Monolith vs microservices: which should you choose?
Start with a well-structured monolith unless you have several teams and clear service boundaries, because microservices add distributed-system complexity such as network failures, data consistency and operational overhead.