Skip to content

Debugging and incident reasoning in a Senior backend interview

This round hands you a symptom — a latency spike, a stalled consumer, a pod that will not start — and grades whether you reach the mechanism before reaching for a fix. Thinking out loud is the deliverable, not a side effect.

The check comes with every pass, covers all 8 competencies of a Senior backend loop, and takes 10–15 minutes. It reports each one as Ready, Borderline or a Gap for this level. It never predicts a pass.

Where this usually breaks down at Senior

These are the distinctions the questions on this axis are written to separate — our judgement from writing them and having them reviewed, not a measurement of candidates. We have run no study, so nothing here is a statistic and none of it is phrased as one.

  1. 1.

    A hypothesis nothing could refute

    “It is probably the database” becomes a hypothesis only when you say what you would see if it were true and what you would see if it were not. Without that, every observation confirms it and the investigation cannot converge.

  2. 2.

    Correlation read as mechanism

    A restart curing the symptom is equally consistent with a memory leak, a backlog, a stale connection pool and corrupted in-process state. It discriminates between none of them, so it is evidence for none of them.

  3. 3.

    Fixing before bounding the blast radius

    Who is affected, and whether it is still getting worse, come before why. Mitigation and diagnosis are different jobs with different urgencies, and conflating them spends an outage on the interesting question rather than the expensive one.

  4. 4.

    Reading the report instead of the effect

    A green deploy, an exit code of zero, a metric that is flat because nothing is reporting into it. Separating what was measured from what was assumed is most of the skill, and it is exactly what gets planted in these questions.

What a strong answer sounds like: A strong answer states what it expects to see if the hypothesis holds, and what would rule it out, before it looks.

How the check measures it

3 questions drawn on this axis alone — 2 medium and 1 hard for a Senior target, because the bar is the level. Hard counts for one and a half times a medium, which is what makes the verdict relative:

Medium rightHard rightVerdict
2 of 21 of 1Ready
1 of 21 of 1Ready
2 of 20 of 1Borderline
0 of 21 of 1Borderline
1 of 20 of 1Gap
0 of 20 of 1Gap

Three questions is coarse and the report says so beside every verdict. It is enough to separate this axis from the other 7, which is what the check is for — there is no total, no percentage and no single number anywhere in the report.

What there is to practise

316 questions on this axis
203 medium, 113 hard · 11 runnable Go exercises. Easy is not drawn for a Senior target.

Every answer cites the source it was checked against.

Drawn from AWS (36), Observability (34), Databases & SQL (33), Messaging & Event Streaming (29), System Design (27).

Practised in the Technical deep dive and Operations and reliability rounds
How things actually work and why they fail: networking, concurrency, performance, security, and diagnosing a symptom to its mechanism. Running it in production: reliability, observability, Kubernetes, infrastructure as code, incidents and cost — graded explicitly since 2026.

Practice is arranged by interview round rather than by topic, because loops are graded that way.

Questions here most often also test reliability & failure modes, operations & sre, data modelling & storage and observability. Nothing on this axis is asked in isolation, which is why the check reports every competency separately rather than averaging them.

One of them, in full

A real question from the free demo, not written for this page. A demo quiz opens with it; the answer, the explanation and the source come when you submit — no account needed.

Debugging and incident reasoninghard

Your team attaches an existing Lambda function to a VPC so it can query an RDS database, selecting the subnets the RDS instance sits in. After the change, the function starts timing out when calling a public third-party API it previously called successfully. Those subnets route 0.0.0.0/0 to an Internet Gateway, the function's security group allows all outbound traffic, and the calls fail on every invocation rather than only on the first after an idle period. What is the cause?

Answer it on the demo

One axis is not the loop

A Senior backend loop tests 8 competencies, and being strong here says nothing about the other 7. The check covers all of them in one sitting and reports each separately, so what comes back is a profile rather than a score.