Skip to content
Gun.io
September 28, 2026 · 13 min read

System Design Interview Questions to Ask Senior Engineers

We spend a lot of time evaluating engineers at Gun.io, and I think one of the easiest mistakes in technical hiring is to confuse familiarity with an interview format for familiarity with the work. Someone can learn how to draw a plausible architecture without having much experience deciding what should actually be built. A system design interview is useful when it gives you a way to examine that distinction.

In practice, an engineer rarely receives a perfectly specified problem. An existing system, incomplete requirements, a budget, a deadline, and a collection of constraints may not become obvious until the work starts. I want to understand how someone makes progress in that environment: which questions they ask, which assumptions they make, what they decide, and what would cause them to change their mind.

The system design interview questions below are a starting point for that conversation. They include technical answer notes, but those notes are not a single approved architecture. A candidate who proposes a different design and explains it well may be giving you a better answer than someone who reproduces every component on your checklist.

What Should a System Design Interview Test

Start with the responsibility you expect the engineer to carry. Designing an internal reporting service for a small team is a different job from designing a payments platform that operates across regions. Both involve system design, but the relevant risks and the appropriate amount of complexity are different.

I would look for five things: whether the candidate clarifies requirements, uses scale estimates meaningfully, protects data correctness, considers failure and recovery, and explains decisions clearly. The important part is how those things connect. An estimate should affect a design decision. A reliability requirement should affect cost. A decision about consistency should reflect what the product can tolerate.

Someone who recognizes that a simple design meets the requirements may be showing excellent judgment. Do not reward unnecessary infrastructure just because it makes the diagram look senior.

15 System Design Interview Questions With Strong Answer Notes

Choose one main scenario and use the others for follow-ups or later rounds. The groups below separate questions about scale, data correctness, and recovery so you can emphasize the responsibility the role actually carries.

Scalability and Performance Questions

  1. How would you design a URL shortener with heavy read traffic?

Ask the candidate to establish the read and write volume, expected retention, and latency requirements before choosing storage. They might use a counter encoded in base62, preallocated key ranges, or random identifiers with collision handling. Each choice creates different coordination and predictability concerns.

Caching may help, particularly when a few links account for much of the traffic. Ask what happens when one link becomes unusually popular, and how redirects affect the analytics requirement. A permanent redirect can be cached by clients, which may reduce the number of visits observed by the service. The useful signal is whether the candidate connects those details to what you asked the system to do.

  1. How would you design a news feed for millions of users?

The candidate should ask about freshness, ranking, follower distribution, and whether every user needs the same product behavior. Precomputing feeds when someone posts can make reads cheaper, but accounts with very large audiences make writes expensive. Assembling feeds when someone reads moves that cost in the other direction.

A hybrid can be reasonable, but it should follow from the workload rather than appear as a memorized answer. Ask about pagination, deleted posts, privacy changes, and ranking updates. Those questions reveal whether the proposed feed remains correct after the initial happy path.

  1. How would you design a rate limiter for a public API?

A rate limiter needs a policy before it needs an algorithm. Are limits applied to users, tenants, API keys, IP addresses, or some combination? Are brief bursts acceptable? What happens if the limiter becomes unavailable?

Fixed windows are simple, but permit bursts across window boundaries. Sliding logs offer precision at a greater storage cost. Sliding counters approximate a rolling window, while token buckets allow a configured amount of bursting. Ask the candidate to select an approach and explain how concurrent requests update shared state safely. Returning a useful throttling response is part of the design, as is deciding whether a failure should permit or block traffic.

  1. How would you design search for a large product catalog?

Look for a distinction between the authoritative product data and the index used to search it. The candidate should explain how changes reach the index, how failed updates are detected, and how a reindex happens while users continue searching.

Then bring the conversation back to the product. What happens when a result contains an item that has sold out? Where is inventory checked before purchase? How do filters, typo tolerance, and ranking affect the design? A latency target is useful when it comes from a requirement and has a credible measurement plan; an impressive number chosen without context tells you very little.

  1. How would you design video upload and processing?

Large uploads and long processing jobs make this a useful question about asynchronous work. Direct uploads to object storage, resumable transfers, queued processing, and status tracking are plausible components. The candidate should explain how authorization works and how processing can be retried without creating inconsistent output.

Ask about partial uploads, a worker dying halfway through a job, and the cost of generating multiple resolutions. A CDN and adaptive streaming may fit the delivery requirement. What matters is whether the candidate can distinguish what every upload needs from work that can be deferred or avoided.

Data and Consistency Questions

  1. How would you prevent double booking?

Ask what is being reserved. A fixed seat at a fixed time can sometimes be protected with a unique constraint. Overlapping appointment intervals require a different constraint or concurrency strategy. Locking, transactions, and optimistic checks should be discussed in relation to that actual model.

Then introduce checkout. A temporary hold expires while payment is in progress; what happens next? The candidate needs to define which system owns the final reservation, how retries behave, and how to handle a payment that succeeds after inventory is released. This is where an apparently straightforward database question becomes an ownership question.

  1. How would you design a payment workflow that handles retries?

Look for idempotency keys, explicit payment states, and reconciliation with the payment provider. The candidate should distinguish a confirmed failure from an unknown outcome. A timeout does not tell you whether money moved.

Ask how the same operation is identified across retries, what happens when a key is reused with different parameters, and how an ambiguous result is resolved through provider-supported idempotency or status checks. An outbox can keep a local state change and its event consistent, but it does not make an external payment atomic with your database. A good answer respects that boundary.

  1. How would you keep a cache and database consistent?

First establish what stale data would mean. An old product description and an old available balance create different risks, and they should not receive the same caching policy.

Cache-aside with invalidation is common, but concurrent reads and writes can repopulate stale values. Time-to-live settings limit some exposure without eliminating the race. Write-through or change-driven invalidation introduces different coordination and failure concerns. Ask the candidate to describe the failure sequence, not simply name a pattern. For correctness-sensitive decisions, they should explain when to consult authoritative state rather than trust a cache.

  1. How would you migrate a large database while keeping the application available?

A credible migration plan covers compatibility, backfill, validation, cutover, and recovery. Expanding a schema before removing the old path is often useful. Moving between datastores may require change capture or carefully controlled dual writes; neither should be treated as effortless synchronization.

Ask how the candidate detects missed updates, limits load from the backfill, and proves that the new system represents the same business data. Row counts alone are insufficient. Then ask about rollback after writes have moved to the new system. If those writes cannot reach the old one, switching a flag back may create a second problem instead of solving the first.

  1. How would you design an audit log for sensitive changes?

An audit record should make it possible to understand who acted, what they did, when it happened, and which resource was affected. Authorization context and correlation identifiers may also matter. Ask what belongs in the record and what should be excluded because it contains credentials or unnecessarily sensitive information.

Append-only application behavior is helpful, but an administrator with storage access may still alter records. The candidate should explain the required level of tamper resistance, access separation, retention, and investigation support. Hash chaining can help detect some alterations; it needs a trusted reference and does not, by itself, solve every integrity problem.

Reliability and Security Questions

  1. How would you design chat that recovers from disconnections?

Ask the candidate to separate message acceptance, persistence, delivery, and reading. A WebSocket connection can support live delivery, but a connection alone does not establish that a message was stored or received.

Conversation sequence numbers and resumable history can help clients recover missed messages. The design also needs to address duplicates, ordering, and authorization when someone reconnects. Presence indicators can tolerate uncertainty in ways that message history may not. I would want the candidate to make those distinctions explicit rather than promise that everything is real time and perfectly consistent.

  1. How would you avoid duplicate notifications?

Retries and at-least-once delivery make duplicate handling an ordinary design concern. A notification identifier and a record of processing are useful, but ask what happens if the provider accepts a message and the worker crashes before recording success.

That gap may require provider idempotency or a deliberate trade-off between possible duplicates and possible omissions. A database check before sending does not close it automatically. Then discuss user preferences, quiet hours, and separate capacity for urgent notifications. A password reset should not inherit the delivery behavior of a marketing campaign merely because both send email.

  1. How would you handle a regional outage?

Start with recovery time and acceptable data loss. Backup and restore, pilot light, warm standby, and active-active operation offer different cost and operational profiles. Their actual recovery performance depends on the implementation and the team’s ability to execute it.

Ask about replication lag, conflicting writes, failover testing, and the return to normal operation. Active-active does not automatically mean zero downtime or zero data loss. Replication can also propagate deletion or corruption, so it does not replace recoverable backups. The right answer is the least complicated approach that credibly meets the stated requirement.

  1. How would you enforce access control in a multi-tenant application?

The design should make tenant context explicit and enforce authorization at the relevant boundaries. Scoped data access, database row-level security, or separate databases may help, depending on the application. None removes the need to examine background jobs, caches, exports, files, and administrative tools.

Ask how cross-tenant access is tested, including direct requests with another tenant’s resource identifier. Role-based and attribute-based policies can both be appropriate. A strong candidate explains where the policy is enforced and how a future engineer avoids accidentally bypassing it.

  1. How would you investigate recurring failures in a distributed system?

Look for a method that links user-visible symptoms to evidence: metrics, logs, traces, recent changes, and the behavior of dependencies. Ask how the candidate prioritizes restoring service while preserving enough information to understand the failure.

Timeouts, retries, backoff, circuit breakers, and isolation can help, but inappropriate retries can increase load on a struggling dependency. Ask for a concrete incident and what changed afterward. A useful incident review produces actions with owners and a way to confirm they worked. The name of the review process is less interesting than whether the same failure becomes less likely.

How Should You Run the Design Session

Set the Scenario Constraints and Time Limit

Use one main scenario and leave time to examine it properly. A 60-minute session might allocate ten minutes to requirements, twenty to the initial design, twenty to a deeper discussion, and ten to questions. Adjust that structure to the job, and give candidates comparable information and assistance.

Let the Candidate Clarify Requirements and Explain Trade-Offs

Give the candidate room to ask questions before drawing the architecture. Answer factual questions directly, and ask them to explain choices that depend on judgment. I want to see which constraints they establish independently and which assumptions they make visible. A candidate can reach a different design from yours and still be reasoning well.

Use a Whiteboard Only When It Helps Communication

A shared diagram helps when it clarifies the discussion. A written outline can serve the same purpose. Let the candidate explain the data flow and failure behavior in a form you can both follow; drawing skill should not become an accidental hiring criterion.

How Do You Score Architectural Judgment

A Consistent Rubric for Scale Failure Handling and Communication

Record the evidence independently before discussing the candidate with other interviewers. A useful note explains which estimate affected the design or how the candidate handled an uncertain payment outcome. The score should make those observations easier to compare.

DimensionEvidence to record
RequirementsWhich constraints the candidate surfaced and which assumptions they stated
ScaleWhether estimates affected choices and exposed a bottleneck
CorrectnessHow concurrency, retries, and unknown outcomes were handled
RecoveryHow failures would be detected, contained, and reversed
JudgmentWhy an option was chosen and when it should change
CommunicationWhether another engineer or stakeholder could act on the explanation

A simple four-point scale can organize the notes, from unable to apply the concept through independent reasoning about the wider system. It is an interview aid, not a validated predictor of performance. Write the evidence first and use the score to summarize it.

Distinguish Senior Ownership From Staff Level Scope

A senior role may require independent ownership of a substantial service. A staff role may add decisions across teams, shared infrastructure, and a migration path that other groups can follow. Define the expected responsibility before the interview. Applying staff expectations to a senior role can reject someone well suited to the work.

Common Mistakes Interviewers Make

Keep follow-ups neutral. Asking where latency comes from reveals more than suggesting a cache. Record what the candidate identified independently and where you supplied a hint. Avoid grading against one preferred architecture, testing unrelated trivia, or spending the whole session on one component while leaving correctness and recovery unexplored.

How Gun.io Assesses System Design in Its Vetting

A design conversation gives you a sample of reasoning. Work history, implementation ability, references where appropriate, and performance in the actual environment provide additional evidence. I would be careful about treating a polished interview as proof of dependable delivery, or a hesitant presentation as proof that an engineer cannot do the work.

At Gun.io, the question behind evaluating an engineer is practical: what responsibility can we reasonably trust this person to carry in a particular engagement? System design is useful when it helps answer that question. It becomes much less useful when it turns into a competition to produce the most elaborate architecture in an hour.

Frequently asked questions

What are good system design interview questions?

Use scenarios that resemble the role: reservations for concurrency, payments for retries and uncertain outcomes, migrations for operational planning, or multi-tenant applications for access control. Choose the scenario because its decisions matter to the work.

How Do You Evaluate a System Design Interview?

Look at the requirements they establish, the assumptions they make, and how they defend and revise decisions. Examine correctness and recovery alongside scale, and record specific evidence using a consistent rubric.

What Separates a Staff Engineer System Design Answer From a Senior Engineer Answer?

A senior role often involves independent ownership of a substantial system or feature area. A staff role may require decisions across teams and systems. Those labels vary, so define the expected responsibility before applying the title.

Should every engineer receive a system design interview?

Only when design responsibility is material to the role. The interview should reflect the work you need done, and implementation or code review may be more informative for a narrower position.

How Long Should a System Design Interview Be?

A 45- to 60-minute session can give you time for requirements, a design, and meaningful follow-ups. Use fewer scenarios in greater depth, and adjust the time to the responsibility you need to assess.

Should Candidates Use a Whiteboard in a System Design Interview?

Use a whiteboard or shared diagram when it helps both people follow the design. A clear written explanation is also useful. Assess the reasoning and communication rather than the visual polish.

Hire for Decisions That Hold Up in Production

I would use the design interview to understand which decisions an engineer can carry independently, then check that picture against implementation and work history. If you are evaluating a hire with Gun.io, start with the system, its constraints, and the responsibility you need covered. That gives us a useful basis for discussing fit.

Gun.io

Sign up for our newsletter to keep in touch!

This field is for validation purposes and should be left unchanged.

© 2026 Gun.io