Skip to content

System design & trade-offs in a Staff backend interview

The Staff version of the same round changes what is being graded. The system is usually one that already exists, with traffic on it and a team around it, and the interesting part is the second design — the one you did not choose, and why.

The check comes with every pass, covers all 10 competencies of a Staff backend loop, and takes 10–15 minutes. It reports each one as Ready, Borderline or a Gap for this level. It never predicts a pass.

Where this usually breaks down at Staff

These are the distinctions the questions on this axis are written to separate — our judgement from writing them and having them reviewed, not a measurement of candidates. We have run no study, so nothing here is a statistic and none of it is phrased as one.

  1. 1.

    The trade-off is stated but not priced

    Naming a trade-off is table stakes above Senior. What separates a Staff answer is saying what it costs and who pays: latency on which path, spend, operational burden on which team, and over what horizon.

  2. 2.

    One design, defended

    A single architecture with no alternative reads as a preference rather than a decision. The expectation is two workable designs and a reason for the choice that somebody else could re-derive from the same constraints.

  3. 3.

    Scale asserted rather than bounded

    “This scales” is not a claim anyone can check. The useful answer says where it stops working — which resource saturates first, at roughly what load — and what the next move is when it does.

  4. 4.

    No migration

    The design is for a system that exists and cannot be turned off. How traffic moves from the current shape to the proposed one, what runs in both meanwhile, and which steps are not reversible is the Staff half of the question, and it is where most answers simply end.

What a strong answer sounds like: A strong answer names the condition under which it would pick the other design, and what switching would cost then rather than now.

How the check measures it

3 questions drawn on this axis alone — 1 medium and 2 hard for a Staff target, because the bar is the level. Hard counts for one and a half times a medium, which is what makes the verdict relative:

Medium rightHard rightVerdict
1 of 12 of 2Ready
0 of 12 of 2Ready
1 of 11 of 2Borderline
0 of 11 of 2Gap
1 of 10 of 2Gap
0 of 10 of 2Gap

Three questions is coarse and the report says so beside every verdict. It is enough to separate this axis from the other 9, which is what the check is for — there is no total, no percentage and no single number anywhere in the report.

This axis is weighted 1.5× when the report picks your highest-risk gap, because loops at this level weight it more than the rest.

What there is to practise

168 questions on this axis
132 medium, 36 hard. Easy is not drawn for a Staff target.

Every answer cites the source it was checked against.

Drawn from System Design (77), Messaging & Event Streaming (27), Networking & APIs (14), Databases & SQL (13), Google Cloud (10).

Practised in the System design round
Design a system to requirements and defend the trade-offs: storage, messaging, caching, capacity and cost. The round weighted most in 2026 loops.

Practice is arranged by interview round rather than by topic, because loops are graded that way.

Questions here most often also test technical judgment: decisions and trade-offs, distributed systems, messaging & async integration and data modelling & storage. Nothing on this axis is asked in isolation, which is why the check reports every competency separately rather than averaging them.

One of them, in full

A real question from the free demo, not written for this page. A demo quiz opens with it; the answer, the explanation and the source come when you submit — no account needed.

System design & trade-offsmedium

A team removes a deprecated int32 legacy_id = 4; field from a protobuf message and ships the new schema. Six months later someone adds string region = 4;. Old clients still exist in the field. What goes wrong, and what should have been done when the field was removed?

Answer it on the demo

One axis is not the loop

A Staff backend loop tests 10 competencies, and being strong here says nothing about the other 9. The check covers all of them in one sitting and reports each separately, so what comes back is a profile rather than a score.