Home / Blog

Technical Interview Follow-Up Questions: The 4-Rung Ladder

The follow-up ladder for technical interviews (base, perturbation, derivation, ownership), with worked frontend, backend and behavioral examples.

In a nutshell

  1. Any single question with a known answer is now a lookup. A hidden assistant can fetch it in the time it takes to say "great question". So the first question is the least informative part of an interview.
  2. The signal lives in the follow-ups. I use a four-rung ladder: base (a scoped problem), perturbation (change one constraint on screen), derivation (why this and not the obvious alternative, and how would you know it's wrong), and ownership (when did you last do this, and what broke).
  3. Each rung removes something an assistant depends on: the pre-generated answer, the generic answer, and finally the answer that doesn't need a life behind it.
  4. Score the rungs, not the vibe. A perfect base answer on its own should never score above a 2 out of 4.
  5. Below are three worked ladders (frontend, backend, behavioral), scoring anchors and a cheat sheet.

Why do follow-up questions matter more than the first question?

Because the first question is the one everybody can prepare for, and the one a tool can answer.

Canva's engineering team found that AI "can trivially solve traditional coding interview questions". Fabric, which sells AI interviewing with cheating detection, reports that among interviews its detector flagged, 45% involved overlay tools that read the screen and 34% voice-mode assistants that hear the question. Vendor data, but the mechanics are clear: a fixed question goes in, an answer comes out.

Follow-ups depend on what the candidate just said, change while they're talking, and at the deepest level depend on their own past. No tool has that context.

There's an older reason too. Structured interviews, with planned questions and anchored scoring, are the top-ranked selection method in Sackett and colleagues' 2022 re-analysis, with a validity of about .42 against .19 for unstructured interviews. Planned follow-ups are how you get structure without asking robotic questions.

Yet Karat's survey of 400 engineering leaders found only 27% are prioritising training interviewers to measure AI skills. Karat sells interviewing services, so treat that as vendor data. If your interviewers improvise follow-ups, planned ladders are the cheapest upgrade you have.

What are the four rungs?

Rung What you do What it measures What it takes from an assistant
0. Base Ask a realistic, scoped problem Can they start? Baseline approach Nothing. Assume it can be fully assisted
1. Perturbation Change one constraint, input, scale or requirement, on screen Can they adapt their own solution? The pre-generated answer
2. Derivation "Why this and not X?" "How would you know it's wrong?" Do they understand the reasoning? The generic answer
3. Ownership "When did you last hit this? What shipped? What broke?" Is the skill theirs? The answer with no history behind it

A full ladder takes 12–20 minutes. In a 45-minute round, run two ladders, not five questions.

How do I run a ladder well?

Four rules.

Show it on screen, don't just say it. Put the problem in a doc or editor you control and reveal it in parts. Audio pipelines hear everything you say. A constraint typed on screen mid-session forces a recapture. It's also fairer to candidates working in a second language.

Change details live. Prepare 3–4 perturbations per question. Change one thing at a time: input (empty, huge, duplicated, out of order), scale (×100), requirement (undo, audit, offline), constraint (no new dependencies, 200 ms budget), or failure (a dependency times out 5% of the time).

Ask for brute force first.

"Start with the simplest thing that works, even if it's slow. We'll improve it together."

It lowers the stakes for nervous candidates, and it makes a sudden jump to a polished, optimal answer conspicuous.

Set think-aloud norms at the start.

"I care more about how you think than whether you finish. Talk me through it, including ideas you reject. Silence is fine, just tell me you need a minute. And asking clarifying questions counts in your favour here."

Clarifying questions are signal. Generated answers rarely include them.

Worked example 1: what does a frontend ladder look like?

Role: mid-level frontend engineer. Time: about 20 minutes.

Rung 0: Base

On screen: a product card with a heart icon and a stub comment, POST /api/favourites/:id.

"Make the heart toggle a favourite. Simplest version that works."

Listen for: it works, and they mention what the user sees while the request is in flight.

Rung 1: Perturbation

Type into the problem doc: "The API takes 800 ms and fails 3% of the time. Users double-click."

"Same component, new reality. What does the user experience now, and what would you change?"

Strong: "Waiting 800 ms feels broken, so I'd update the heart immediately and roll back if the request fails, with a small error message. Double-clicks could send two requests that cancel each other out, so I'd ignore clicks while one is pending, or send the desired final state rather than 'toggle'." Then they edit their own handler.

Weak: a general explanation of optimistic UI that never mentions rollback or the double-click, or a full rewrite using a data-fetching library that wasn't there a minute ago.

Rung 2: Derivation

"You chose to send the final state instead of a toggle. Why? And if we shipped this and some users' favourites were silently wrong, how would you find out?"

Strong: "A toggle isn't idempotent. If a retry lands twice, you're back where you started." On detection: one concrete test, such as mocking two requests resolving in reverse order, or one metric, such as comparing client-side and server-side favourite counts for a sample of users.

Weak: a balanced essay on tradeoffs with no commitment, or a list of monitoring tools with no first step.

Rung 3: Ownership

"When did you last build something optimistic like this? What was the product, and what went wrong after it shipped? How did you find out?"

Listen for (non-technical): do they mention what happens when the request fails or the user clicks twice? Do they describe a real product, with a real problem, that gets more specific when pushed?

Worked example 2: what does a backend ladder look like?

Role: mid-level or senior backend engineer. Time: about 20 minutes.

Rung 0: Base

On screen: an endpoint stub, POST /imports/customers, accepting a CSV with email, name, plan.

"Write the handler that imports these rows into the customers table. Simplest version first."

Listen for: they ask what happens if an email already exists.

Rung 1: Perturbation

Type: "Files can be 2 million rows. About 5% of rows are malformed. The request times out at 30 seconds."

"What happens now?"

Strong: "Loading 2 million rows in the request will time out and use a lot of memory. I'd accept the upload, store the file, return a job ID, and process it in the background in batches. For malformed rows, I wouldn't fail the whole file. I'd skip them and write them to an error report the user can download."

Weak: correct general advice about background jobs that isn't applied to the code on screen, or a solution that still fails the entire import on the first bad row.

Rung 2: Derivation

"Why batches of a few thousand rather than one row at a time, or the whole file in one transaction? And a customer says 'I uploaded 50,000 rows and only 48,000 showed up.' How do you answer them?"

Strong: a clear reason for the batch size (round trips versus lock duration and memory). On the complaint: "I'd reconcile counts. Rows read, rows inserted, rows updated, rows rejected. They should add up to 50,000. If they do, the error report explains the gap. If they don't, we have a bug, and I'd check whether a failed batch was retried."

Weak: "it depends" with no number, or a list of tools without the reconciliation idea.

Rung 3: Ownership

"Tell me about the largest data import or migration you've run. What went wrong, how did you find out, and how did you clean up?"

Follow-ups: "How many rows?" "How long did it take?" "What did you add so it couldn't happen again?"

Listen for (non-technical): do they talk about not losing rows silently, and how they'd prove the numbers add up? Do they describe a real incident with a number and a fix?

Worked example 3: what does a behavioral ladder look like?

Any model produces a tidy STAR story on demand, so in a behavioral round the base answer is worth almost nothing and the ladder is everything.

Rung 0: Base

"Tell me about something you shipped that broke in production."

Rung 1: Perturbation (change the story's conditions)

"Suppose nobody noticed for a week instead of a day. What would have been different?"

"Suppose you'd been on holiday when it broke. Who would have fixed it, and would they have been able to?"

Strong: they reason about their actual team and systems. "A week would have meant corrupted invoices, because the nightly job built on the bad data."

Weak: a generic principle that could apply to anyone. "I'd make sure we had good monitoring and documentation."

Rung 2: Derivation

"What made you sure of the cause at the time? What would have changed your mind?"

Strong: specific evidence, including weak evidence they admit was weak. Real incidents usually have an uncomfortable part: a wrong first guess, a rollback that didn't help.

Rung 3: Ownership (the artifacts)

"Where would I find a trace of this today? A postmortem, a ticket, a PR? Who else would remember it, and what would they say you got wrong?"

Strong: they name an artifact and a person, and give the other side a fair hearing.

Weak: the story is perfectly structured, nobody made a mistake, and the outcome is a round-number improvement with no cost.

Listen for (non-technical): does the story get more specific or more polished when you push? Can they say what a colleague would say they got wrong?

How do I score a ladder?

Write one line of evidence per rung before the debrief. "Rung 1: edited own handler, found the double-click issue unprompted" is evidence. "Seemed solid" isn't. Then give the whole ladder one score:

Score What it looks like
1 Stuck at Rung 0, or the answer doesn't change when the problem does
2 Solid Rung 0. Perturbations get the original answer with new nouns. Rung 2 is vague
3 Adapts correctly at Rung 1. Names a concrete check at Rung 2. Rung 3 has specifics
4 All of 3, plus raises a failure mode you didn't ask about, or says "I don't know, here's exactly how I'd find out"

A perfect Rung 0 caps at 2. The score lives where hidden help runs out.

What's the difference between a strong answer and a scripted one?

These are patterns to probe, never proof. Each has an innocent explanation.

Pattern Innocent explanation What to do
Answers the old problem after a perturbation Didn't notice the change Point at the change, ask again
Balanced lists with no commitment Cautious personality "Pick one. Which first?"
Restarts from scratch instead of editing Prefers clean slates "Can you get there by changing your existing code?"
Generic story that doesn't sharpen NDA, poor memory, long ago Ask for a different example
Never asks a clarifying question Anxious about seeming slow Invite questions explicitly

You never write "used AI" on a scorecard. You write what they did and didn't demonstrate. Nerves, accents, second-language phrasing and slow internet are not evidence.

What if the candidate is junior and has no history to own?

Rung 3 doesn't need a job. It needs something they did: a capstone project, an internship, open source, or the take-home they just submitted. "What was the hardest bug in your final-year project, and how did you find it?" works fine. You're testing that their account of their own work gets more specific under questions.

Cheat sheet

  • [ ] Two base questions per 45 minutes, each with 3–4 prepared perturbations
  • [ ] Problem in a doc you control, revealed in parts
  • [ ] Opening script: brute force first, think aloud, questions welcome
  • [ ] Rung 1 typed on screen; Rung 2 asks "why not Y?"; Rung 3 asks "what shipped, what broke?"
  • [ ] One line of evidence per rung, one 1–4 score per ladder

What should I do next?

Pick one question from your current loop and write three perturbations for it. That alone changes what your next interview measures.

If you want to check the rest of your process, the free cheat-risk audit shows where you're exposed, and the free AI interview policy template sets the rules candidates follow. For where ladders fit in a full loop, read should you allow AI in coding interviews, and for recruiters using a simplified version in first screens, how to screen developers as a non-technical recruiter.

Unscripted, the paid kit has ready-made ladders like these for each major stack and role, with listen-for cues and scoring anchors, if you'd rather not write them from scratch.

Hrishikesh Pardeshi is co-founder & CTO of Flexiple and has spent 10 years hiring and vetting developers. He wrote Unscripted, a vendor-neutral kit for running developer interviews in the AI era.