Skip to content

Technical judgment: decisions and trade-offs in a Staff backend interview

This is the axis every ladder above Senior names, and it is usually graded in conversation rather than on a whiteboard: a decision you actually made, the options you weighed, and what you gave up. The interviewer is testing whether the reasoning survives being re-run with different inputs.

The check comes with every pass, covers all 10 competencies of a Staff backend loop, and takes 10–15 minutes. It reports each one as Ready, Borderline or a Gap for this level. It never predicts a pass.

Where this usually breaks down at Staff

These are the distinctions the questions on this axis are written to separate — our judgement from writing them and having them reviewed, not a measurement of candidates. We have run no study, so nothing here is a statistic and none of it is phrased as one.

  1. 1.

    Reversibility never enters the comparison

    The cost of a decision is its downside multiplied by how hard it is to undo. A cheap reversible choice and a one-way door deserve very different amounts of analysis, and treating them alike is how teams spend a fortnight on the first and an afternoon on the second.

  2. 2.

    Optimising for the technically better system

    A Staff decision is constrained by the team that has to run the result, by the deadline, and by everything else already in flight. A better architecture nobody has the capacity to operate is the worse decision, and saying so out loud is the signal.

  3. 3.

    No condition that would falsify it

    A decision with no stated trigger for revisiting it cannot be reviewed later; it quietly becomes the way things are. “We move to X if write volume passes Y, or if the p99 stops meeting Z” is a decision. “X felt right” is a preference.

  4. 4.

    The do-nothing option unscored

    Keeping the current system, with its known cost, is a real option and is often the correct one. Leaving it out makes every alternative look good by default, and the missing baseline is noticed.

What a strong answer sounds like: A strong answer states what would have changed the decision — and can point at an occasion when something did.

How the check measures it

3 questions drawn on this axis alone — 1 medium and 2 hard for a Staff target, because the bar is the level. Hard counts for one and a half times a medium, which is what makes the verdict relative:

Medium rightHard rightVerdict
1 of 12 of 2Ready
0 of 12 of 2Ready
1 of 11 of 2Borderline
0 of 11 of 2Gap
1 of 10 of 2Gap
0 of 10 of 2Gap

Three questions is coarse and the report says so beside every verdict. It is enough to separate this axis from the other 9, which is what the check is for — there is no total, no percentage and no single number anywhere in the report.

This axis is weighted 1.5× when the report picks your highest-risk gap, because loops at this level weight it more than the rest.

What there is to practise

374 questions on this axis
277 medium, 97 hard. Easy is not drawn for a Staff target.

Every answer cites the source it was checked against.

Drawn from Microsoft Azure (75), System Design (40), Messaging & Event Streaming (37), AWS (36), Databases & SQL (33).

Practised in the Behavioral and judgment round
Decisions with facts on the table: trade-offs, influence, incidents and ownership. Not what you would say — what follows from the situation.

Practice is arranged by interview round rather than by topic, because loops are graded that way.

Questions here most often also test system design & trade-offs, data modelling & storage, security and networking. Nothing on this axis is asked in isolation, which is why the check reports every competency separately rather than averaging them.

One of them, in full

A real question from the free demo, not written for this page. A demo quiz opens with it; the answer, the explanation and the source come when you submit — no account needed.

Technical judgment: decisions and trade-offshard

A Redis cache under allkeys-lru holds about a million keys and serves a working set of roughly 200,000 hot keys at a 97% hit rate. Every night a reconciliation job reads ten million keys once each through the same cache-aside path — every miss is fetched and written into Redis — over twenty minutes. For an hour after the job the hit rate sits near 30% and climbs slowly back. The team proposes allkeys-random, citing the reference's advice for workloads that read keys in a cycle. What is happening, and which policy fits — and with what caveat?

Answer it on the demo

One axis is not the loop

A Staff backend loop tests 10 competencies, and being strong here says nothing about the other 9. The check covers all of them in one sitting and reports each separately, so what comes back is a profile rather than a score.