The AI-era developer interview kit · Edition 1

Your interviews can’t tell you who can code anymore. A webcam won’t fix that.

Candidates can have an AI in their ear that your screen share never shows. Unscripted is a vendor-neutral kit for redesigning developer interviews so they produce real signal anyway. Not with surveillance. With format: explicit AI rules, stages that score judgment, and follow-up questions a pre-generated answer can’t survive.

Solo $79 · Team $199 · 30-day no-questions refund · compare plans
4Books
12Playbook chapters
164Laddered questions
11Work samples
12Templates
7Scorecard sheets
01 — What changed?

The interview stayed the same. The candidate’s toolkit didn’t.

A remote technical interview used to be a reasonable test. You shared a problem, watched someone solve it, asked a few questions.

Now there is software built to sit inside that call. Overlay assistants read the screen and stay invisible to screen share. Voice-mode models hear your question and draft an answer while the candidate says “good question.” Further along: proxies who interview on someone else’s behalf, and the occasional face-swapped video.

The best-known example, Interview Coder, was built to stay hidden from screen share, and its creator recorded himself using it in an Amazon interview (source). It became Cluely, which sells a tier marketed as hidden from meeting screen-sharing software (source).

So a correct answer on a shared screen is no longer evidence that the person on the call can produce it. That is the whole problem, and it’s a format problem.

71%

of 400 senior engineering leaders surveyed by Karat say AI makes technical skills harder to assess.

Karat, 2025. Vendor data: Karat sells interviewing services. Source
20%

of US professionals in a Blind survey said they had secretly used AI during a job interview (503 of 2,510).

Blind, April 2025. Self-selected sample. Source
6%

of 3,000 job candidates admitted posing as someone else in an interview, or having someone else pose as them.

Gartner candidate survey, 2025. Source
48%

of technical-role interviews on Fabric’s platform were flagged by its cheating detector, against 12% for sales roles (19,368 interviews).

Fabric, 2026. Vendor data; sample mostly India-based. Source

Four numbers, four different methods, and none of them is a census. Read together, they point the same way: assume some of your candidates have help, and design the interview so that it doesn’t matter much.

02 — The obvious fixes

Why doesn’t a webcam, proctoring or a detector fix it?

Because each one inspects the wrong layer. And each one makes honest candidates look suspicious while a prepared one routes around it.

Sees the face, not the help

Webcams

The assistant runs on a second device, or in a window the share doesn’t capture. The camera shows a person glancing slightly off-screen.

That is also what nervous, neurodivergent and second-language candidates look like. So you get false positives and no proof.

Watches the browser. The help isn’t there.

Proctoring software

Tab and focus tracking catches someone switching windows. Among the interviews Fabric’s detector flagged, Fabric reports that 79–82% used methods tab-switch proctoring can’t see, mainly overlay tools and voice-mode models (source, vendor data).

An arms race the candidate can keep winning

AI detectors

Detectors work by inspecting the candidate’s machine: windows, mic access, network requests (source). The hiding tools update, then the detectors update.

Meanwhile you’re asking a stranger to install monitoring software, and you’d have to defend a probability score in a rejection.

Detection is an arms race candidates win. Format redesign wins.

This isn’t a fringe view. Canva now expects candidates to use AI tools in technical interviews, after its own tests showed AI “can trivially solve traditional coding interview questions” (source). Meta is piloting a coding interview with an AI assistant built in (source).

Google went the other way: Sundar Pichai said it will add at least one in-person round “just to make sure the fundamentals are there” (source). Opposite choices, same move. Both are format decisions, not detection decisions. Most teams can’t fly every candidate in, so the kit shows you how to get that signal remotely.

03 — The method

So what does work? Decide what each stage measures. Then make every question a ladder.

Every stage goes in one of two lanes. Neither lane needs a camera, a browser extension or an accusation.

Lane A · AI-allowed

Measures judgment around the tool

Any stage you don’t supervise live is in this lane, whether you say so or not. So say so, and score what matters.

  • Did they verify what the AI produced?
  • Did they catch the planted bug, the misleading requirement, the security trap?
  • Are the trade-offs explicit and defensible?
  • Speed and product sense with real tools
Lane B · Independent

Measures reasoning when the tool is wrong

Enforced by format: live, conversational, laddered and tied to the candidate’s own history. Not by cameras.

  • Baseline reasoning on a problem that keeps moving
  • Debugging their own code, live
  • Explaining decisions made in this room
  • Ownership of things they actually shipped
  1. Decide what each stage measures

    One or two competencies per stage, written down, with a lane. If a stage can’t be described in one line, merge it or drop it.

  2. Make the AI policy explicit

    Written, sent with the invite, repeated at the start of each stage. In one vendor’s data, an upfront honesty contract cut unauthorised help on assessments from 28% to 13% (Talogy).

  3. Allow AI where the job does

    Paid, 2–4 hour work samples with three planted issues. Score verification and trade-offs, then run a live oral defense where they change their own code.

  4. Ladder every question

    Base question, then perturb it, then make them derive it, then tie it to something they shipped. Try one below.

  5. Verify identity proportionally

    Tiered checks, applied to everyone at a tier, never to people who “seem off.” Inconsistency is a signal. Accents and nationality aren’t.

  6. Validate in the first 30 days

    Turn the claims from your loop into week-by-week deliverables, so by day 30 each claim is confirmed or you know exactly which one didn’t hold.

What do I ask after they answer? · One ladder from the Playbook · Backend, 20 min Rung 0 of 3
shared-doc · you control this
// Route stub
POST /webhooks/payments

{
  "event_id": "evt_81f2",
  "type": "payment_succeeded",
  "order_id": "ord_5521",
  "amount": 4900
}

You say“Write the handler: mark the order paid and send the receipt email. Simplest version first.”

Strong

Makes it work, and verifies the webhook signature without being asked. A cheap, high-signal habit.

Scripted

Often indistinguishable. A stable question has a retrievable answer. That’s why Rung 0 is only worth 10% of the score.

Listen for Do they check that the request really came from the payment provider?

R0 10%R1 30%R2 30%R3 30%

A pause is not evidence. Nerves, accents, slow connections and neurodivergence all produce pauses. The ladder is built so you never have to guess intent: you score what was demonstrated at each rung, and a pattern is only ever a reason to ask the next one.

04 — What’s inside

Four books, 12 templates and the scorecards to run them.

Every file opens with a short “In a nutshell” summary, so you can hand any single chapter to an interviewer and they can use it the same day.

Book 01 · 12 chapters

The Playbook

The method, end to end: what broke, the two lanes, five reference loops with time costs, the follow-up ladder, AI-allowed work samples, live rounds, identity and proxy fraud, references and the first 30 days, a recruiter’s screen, interviewer training, and the legal and fairness basics.

  1. 00Start Here
  2. 01What Broke: The Threat Model
  3. 02Decide What Each Stage Measures
  4. 03Redesign the Loop
  5. 04The Follow-Up Ladder
  6. 05Run AI-Allowed Work Samples
  7. 06Running Live Interviews in 2026
  8. 07Identity and Proxy Fraud: Proxies, Bait-and-Switch, Deepfakes and Fake Workers

+ 4 more

Book 02 · 164 questions · 12 stacks

The Question Bank

One file per stack or role. Every question ships with what it probes, the problem to show, perturbation, derivation and ownership rungs, strong vs scripted signals, and a “listen for” line for non-engineers. Technical questions add an answer key; senior system design uses full scenario run sheets.

  1. 01Frontend (React + TypeScript) (14 questions)
  2. 02Backend (Node.js) (14 questions)
  3. 03Backend (Python) (14 questions)
  4. 04Backend (Java + Spring Boot) (14 questions)
  5. 05Backend (Go) (14 questions)
  6. 06Full-Stack / Product Engineer (14 questions)
  7. 07Mobile (iOS, Android, React Native) (14 questions)
  8. 08DevOps, SRE and Cloud (14 questions)

+ 4 more

Book 03 · 11 work samples + guide

The Work-Sample Library

AI-allowed take-homes with a scaffold, three planted issues, an anchored 1–4 rubric, a 45-minute oral-defense script with live changes, and rotation variants so solutions don’t leak.

  1. 00How to Run AI-Allowed Work Samples
  2. 01Frontend (React)
  3. 02Backend API
  4. 03Full-Stack Feature
  5. 04Mobile
  6. 05DevOps / Infrastructure
  7. 06Data Engineering
  8. 07ML / LLM Feature

+ 4 more

Book 04 · 12 templates

The Templates

Copy, paste, fill the brackets. Candidate AI policy (three variants), job-post insert, emails, interviewer briefs, suspicion protocol, ID checklist, reference script, debrief, 30-day plan and more.

  1. 01AI Policy for Candidates
  2. 02Job Post Insert (Interview Process and AI Policy)
  3. 03Candidate Emails
  4. 04Interviewer Brief (One Page per Stage)
  5. 05Suspicion Protocol (In-the-Moment Card)
  6. 06Identity Verification Checklist
  7. 07Technical Reference Check Script (20 Minutes)
  8. 08Hiring Debrief

+ 4 more

.xlsx

Scorecard workbook

7 working sheets with the weights and anchors already in place:

  • Loop Designer
  • Stage Scorecard
  • Work Sample Rubric
  • Candidate Comparison
  • Integrity Log
  • Reference Check
  • Funnel Metrics
PDF + Markdown

Read it, or paste it into your wiki

Four PDFs typeset for A4, for print or screen. Every chapter and template also ships as a markdown file that imports cleanly into Notion or any wiki.

Team only

Google Sheets version, team-wide

A one-click Google Sheets copy of the scorecards for your panel, every file as wiki-ready markdown, and a license that covers your whole hiring team.

05 — Free sample

Read one complete question before you buy.

Copied unedited from the Question Bank (Backend (Node.js)). This is the level of detail you get for every question in the kit.

Book 02 · The Question BankBackend (Node.js)

Q1. Why does the import say "done" when nothing is done?

Level: Junior · Format: Debug this snippet · Time: ~8 min

What it probes: Whether they understand how async callbacks behave inside array methods, and what happens to a rejected promise that nothing awaits.

Rung 0: ask

"The client gets { imported: 200 } back right away. But some orders never show up, and the server sometimes crashes a few seconds later. What's wrong?"

app.post('/orders/import', async (req, res) => {
  const { orders } = req.body;

  orders.forEach(async (order) => {
    await db.orders.insert(order);
    await inventory.reserve(order.sku, order.qty);
  });

  res.json({ imported: orders.length });
});

Answer key:

  • forEach ignores the promises returned by the async callback. The handler responds immediately, before any insert finishes.
  • If an insert or reserve rejects, nothing is awaiting it, so it's an unhandled rejection. Since Node 15, the default behaviour for unhandled rejections is to crash the process. That's the "crash a few seconds later."
  • Express 5 forwards rejected handler promises to error middleware, but these rejections happen inside callbacks the handler never awaits, so Express can't see them.
  • All 200 orders also run at once, with no concurrency limit and no transaction. A partial failure leaves orders inserted without reserved inventory.
  • Fix: for...of with await (sequential), or Promise.all / Promise.allSettled with a concurrency limit. Wrap each order's insert and reserve in a transaction. Return per-order results. For big imports, enqueue a job and return 202 Accepted with a job id.

Rung 1: perturb

  • "The client now sends 20,000 orders." Good: don't do this in the request. Use a background job, a progress endpoint, batching, and a concurrency limit.
  • "One bad order shouldn't fail the whole import." Good: Promise.allSettled or per-item try/catch, return which ones failed and why, and make retries idempotent (for example with an upsert on an external order id).

Rung 2: derive

  • "Why not just await Promise.all(orders.map(...))?" (That fixes the ordering, but it fires everything at once, and the first rejection rejects the whole call while the other inserts keep running.)
  • "How would you have caught this before production?" (A test asserting rows exist after the response, the no-misused-promises / no-floating-promises lint rules, and an unhandledRejection log in staging.)
  • "What should happen on unhandledRejection in production?" (Log it with context and let the process crash and restart. Don't swallow it.)

Rung 3: own it

"Tell me about a bug you shipped or debugged that came down to a promise nobody was waiting for. How did it show up?"

Strong signals

  • Spots the forEach problem quickly, and explains the crash too.
  • Raises partial failure and transactions without prompting.
  • Mentions the lint rules that catch floating promises.

Scripted/weak signals

  • "Make the callback non-async" or "add .catch inside" with no fix for the early response.
  • Jumps to Promise.all and doesn't see the concurrency or partial-failure issue when pushed.
  • Can't explain why the process crashes.

Listen for (non-technical recruiter)

  • Good candidates describe it as "the server says it finished before it actually finished."
  • Be wary if they fix only the crash and not the wrong "done" message, or the other way round.
06 — Who it’s for

Who is this for?

People who hire software engineers remotely and have started to doubt what their interviews are telling them.

Founder / CTO · 5–100 people

You’re hiring your first engineers, remotely

One hire who interviewed well and can’t do the job costs you months. You get five reference loops (founding engineer, mid-level, senior/staff, junior, contract) with the time each costs, so you can run a credible loop without a recruiting team.

Engineering manager / tech lead

Your LeetCode round has become a coin flip

You get ladders you can drop into your existing rounds, AI-allowed work samples with planted issues and anchored rubrics, and a 90-minute training session that gets the whole panel scoring the same way.

Recruiter · technical or not

You screen developers but can’t judge code

You don’t need to. Every question has a “listen for” line written for non-engineers, and the Playbook includes a 25-minute recruiter screen and a 25-term glossary with a follow-up question for each term.

Who it isn’t for

  • You want software that flags cheaters automatically. There’s no detector in here, and the kit argues against basing decisions on one.
  • You never talk to candidates live and don’t plan to. The method depends on human follow-up questions.
  • You need legal sign-off. The kit flags where law applies (recording consent, biometric data, automated hiring tools, offer clauses) but it isn’t legal advice.
  • You’re a candidate preparing for interviews. This is written for the other side of the table.
07 — Who wrote it?

Hrishikesh Pardeshi, co-founder & CTO of Flexiple.

I’ve spent ten years hiring and vetting developers for remote teams at Flexiple. This kit is how I think about interviews now that the candidate can bring a second brain to the call.

Most advice on the subject lands in one of two camps: buy a detector, or give up and fly everyone in. I think both miss the point. The question an interview has to answer hasn’t changed: can this person do the work? What changed is which formats can still answer it.

Everything in the kit is something you can run next week with the people you already have. Where a claim needs evidence, it’s cited. Where the evidence is thin, I say so.

  • Co-founder & CTOFlexiple
  • 10 yearsBuilding Flexiple
  • IIM Ahmedabad, NIT TrichyEx-Adobe and Amazon

Details as listed on flexiple.com/about.

08 — Pricing

Pick the license that matches who runs your interviews.

One-time payment. Instant download. Same content in both: the Team license adds the shared formats and covers everyone on your hiring team.

Solo

For one person who runs or screens interviews.

$79 USD · one-time
  • All four books as PDFs: Playbook, Question Bank, Work-Sample Library, Templates
  • Every chapter and template as Notion-importable markdown
  • Scorecard workbook (.xlsx, 7 sheets)
  • License for one person
Buy Solo
30day refund

If the kit isn’t useful, reply to your Gumroad receipt within 30 days of purchase and you’ll get a full refund. No questions, no form. Payment and delivery are handled by Gumroad.

09 — FAQ

Questions people ask before buying

Isn’t this just free blog advice, repackaged?

Some of the ideas exist in pieces: a policy template on one vendor’s blog, a list of red flags on another’s. Most of those pieces end with a pitch for that vendor’s tool.

What you won’t find in one place is the whole loop, vendor-neutral: what each stage measures, the policy, laddered questions with answer keys, work samples with planted issues and rubrics, identity checks, references, the 30-day plan, and the templates to run it. You’re paying for the assembly and the consistency, not a secret.

Does it work for non-technical recruiters?

Yes. That was a design constraint, not an afterthought. Every question has a “listen for” line a non-engineer can use: you don’t judge the code, you judge whether the answer gets more specific or less specific as you push.

The Playbook has a chapter for recruiters: a 25-minute screen, a 0–2 rubric with a “hold” option that sends borderline candidates to an engineer, and a 25-term glossary.

Is this anti-AI?

No. The default policy allows AI in the stages where the job allows it, and scores the judgment around the tool: what they asked, what they kept, what they caught.

The independent stages exist because you also need to see how someone reasons when the tool is wrong. Allowing AI doesn’t lower the bar. It moves it from “can produce code” to “can own code,” which is harder.

Do I still need proctoring or detection software?

The kit doesn’t need any. If you already use a tool, treat its output as one prompt to ask more questions, never as the decision.

Also check the law before adding biometric tools. In Illinois, for example, face-geometry capture falls under BIPA, which requires written consent (source).

What format are the files?

Four PDFs typeset for A4, plus every chapter and template as a separate markdown file that imports into Notion, Confluence or any wiki. Scorecards come as an .xlsx workbook. The Team license adds a one-click Google Sheets copy of the scorecards.

What does the Team license cover?

Everyone who interviews, screens or makes hiring decisions at one company, including contractors helping with your hiring. Put it in your internal wiki and adapt the templates freely.

It doesn’t cover other companies. If you’re a recruiting agency, the Team license covers your own recruiters; you can use the questions with candidates, but you can’t pass the files to clients. Full terms are on the license page.

Is it legal advice?

No. Where law touches interviews (recording consent, biometric data, automated hiring tools, attestation and offer clauses) the kit flags it and tells you to check with counsel. Employment law varies a lot by country and state.

Will this make our interviews longer?

Usually not. The reference loops cost candidates 3–10 hours (the longer ones are paid work samples), cost you 3–9 interviewer hours, and reach a decision in 10 days or fewer.

Inside a round, a ladder replaces five shallow questions with two deep ones in the same 45 minutes.

Is it fair to candidates?

Fairness is built in. Candidates get the rules in writing before they prepare, with a normal way to ask for accommodations. Tells are treated as prompts to probe, never verdicts. Identity checks apply to everyone at a tier. Interviewers write observations, not conclusions.

What if it doesn’t work for us?

Reply to your Gumroad receipt within 30 days and you’ll get a full refund. No questions.

Start this week

Run your next interview on evidence an assistant can’t supply.

Read the Playbook in an evening. Change one stage this week. Send the policy with your next invite.