Skip to content

Reliability & failure modes in a Staff backend interview

At Staff the reliability conversation is less about which timeout to set and more about what the system promises, what that promise costs, and who decided. Expect to be asked what you would let fail.

The check comes with every pass, covers all 10 competencies of a Staff backend loop, and takes 10–15 minutes. It reports each one as Ready, Borderline or a Gap for this level. It never predicts a pass.

Where this usually breaks down at Staff

These are the distinctions the questions on this axis are written to separate — our judgement from writing them and having them reviewed, not a measurement of candidates. We have run no study, so nothing here is a statistic and none of it is phrased as one.

  1. 1.

    A target with no budget

    An availability number nobody is allowed to spend makes every incident equally urgent and nothing prioritisable. An error budget turns the same number into a decision about whether to ship or to stabilise, which is the conversation the role is being hired for.

  2. 2.

    Redundancy that shares a failure

    Two replicas behind one config rollout, one control plane, one certificate or one deploy pipeline are one thing wearing two names. The question is not how many copies exist but what is genuinely independent.

  3. 3.

    A remedy that installs a new failure mode

    Fixes trade. Claiming a unit of work before doing it removes double execution and introduces work that is claimed and never done; adding a retry removes a transient failure and adds load exactly when the system can least take it. Naming what the fix introduces, and what closes that in turn, is the Staff answer.

  4. 4.

    Degradation treated as binary

    Stale, partial or slower is usually better than an error, and what to shed under pressure is a decision taken calmly in advance rather than under load at three in the morning. If nobody has decided, the system decides — badly, and differently each time.

What a strong answer sounds like: A strong answer says what it would shed, in what order, and who agreed to that before the incident.

How the check measures it

3 questions drawn on this axis alone — 1 medium and 2 hard for a Staff target, because the bar is the level. Hard counts for one and a half times a medium, which is what makes the verdict relative:

Medium rightHard rightVerdict
1 of 12 of 2Ready
0 of 12 of 2Ready
1 of 11 of 2Borderline
0 of 11 of 2Gap
1 of 10 of 2Gap
0 of 10 of 2Gap

Three questions is coarse and the report says so beside every verdict. It is enough to separate this axis from the other 9, which is what the check is for — there is no total, no percentage and no single number anywhere in the report.

What there is to practise

379 questions on this axis
240 medium, 139 hard · 5 runnable Go exercises. Easy is not drawn for a Staff target.

Every answer cites the source it was checked against.

Drawn from System Design (57), Python (48), AWS (44), Networking & APIs (43), AI Agents (33).

Practised in the Operations and reliability round
Running it in production: reliability, observability, Kubernetes, infrastructure as code, incidents and cost — graded explicitly since 2026.

Practice is arranged by interview round rather than by topic, because loops are graded that way.

Questions here most often also test debugging and incident reasoning, networking, operations & sre and messaging & async integration. Nothing on this axis is asked in isolation, which is why the check reports every competency separately rather than averaging them.

One of them, in full

A real question from the free demo, not written for this page. A demo quiz opens with it; the answer, the explanation and the source come when you submit — no account needed.

Reliability & failure modeshard

A CPython 3.12 service's watchdog thread must touch a heartbeat file every 100 ms, or a supervisor restarts the process. A nightly export serialised 1.5 million rows with one json.dumps call, 0.9 s, during which the watchdog missed its deadline. The team switched to json.dump(rows, f), expecting streaming to let the watchdog in: it now keeps its deadline, but the export takes 4.2 s, and its window allows 2 s. Which explanation and change are right?

Answer it on the demo

One axis is not the loop

A Staff backend loop tests 10 competencies, and being strong here says nothing about the other 9. The check covers all of them in one sitting and reports each separately, so what comes back is a profile rather than a score.