Home / Blog

Take-Home Assignments and AI: How to Still Get Signal

AI can now do the classic take-home assignment for the candidate. How to design one that still gives signal: planted issues, a usage log, oral defense.

In a nutshell

  1. A take-home you don't supervise is AI-allowed, whatever the instructions say. So stop pretending otherwise and design for it.
  2. A classic take-home ("build a small to-do API") now mostly measures whether the candidate has a decent AI tool. Everyone does.
  3. The fix isn't to drop take-homes. Work samples are still among the better predictors of job performance. The fix is to score judgment, not output.
  4. Four design changes do it: a realistic scaffold repo with three planted issues, an explicit time box and fee, a short AI usage log, and a live oral defense where they change their own code.
  5. The oral defense is the part that makes outsourcing pointless. Someone else can do the take-home. They can't sit the defense for the candidate.

Did AI kill the take-home assignment?

It killed the old kind.

Canva's engineering team tested this directly and wrote that AI "can trivially solve traditional coding interview questions". A self-contained take-home with a clear spec is the easiest version of that. Paste the brief, get the code, tidy it up, submit.

Candidates are open about it, too. In CoderPad's 2025 survey, 12% of developers admitted cheating on a technical assessment at least once, and 47.9% of those used AI to generate or optimise code. CoderPad sells an interview platform, so read that as vendor data. And "cheating" there assumes AI was banned. Many candidates don't see it that way when nobody is watching.

So if your current take-home is a clean spec, a blank repo and a deadline, you're scoring the tool. That's not the candidate's fault. It's the format.

So should I drop take-homes altogether?

No. You'd be throwing away one of the better signals you have.

The most recent large re-analysis of selection research, by Sackett and colleagues in 2022, puts work samples at a validity of about .33, against .19 for unstructured interviews. Structured interviews came top at .42. So a well-run work sample beats a chat by a wide margin, and pairs well with a structured interview.

The key word is "well-run". A work sample predicts performance when it shows how the person works. If AI does the work and nobody checks what the person understood, the validity goes with it.

So keep the take-home. Change what it asks for and what you score.

What does an AI-allowed take-home look like?

Here's the shape I'd use for a mid-level engineer.

Element What it is Why
Scaffold repo 15–40 files in your stack, with conventions, tests that run in one command, and a README AI tools behave differently in an existing codebase than on a blank file. You want the realistic version
Task A small feature that touches existing code Their work has to run through code someone else wrote
Three planted issues A subtle bug, a misleading requirement, a security or performance trap Detection is the strongest single signal of judgment
Time box 2–4 hours, stated in the invite Fair to people with jobs and families. Comparable between candidates
Fee Flat, paid on submission, whatever the outcome Widens your pool and lets you ask for real work
Hand-in Code with commit history, short README, AI usage log, 5-minute walkthrough Gives you a map for the defense
Oral defense 45–60 minutes, live Proves the work is theirs

To scope it, have one of your engineers do the task with AI tools and time them. Multiply by 1.5 to 2. If that's over four hours, cut scope. Don't extend the box.

Then say exactly what to do when time runs out:

"Aim to spend about 3 hours. If you hit 3 hours, stop, commit what you have, and write a short note on what you'd do next. A thoughtful partial solution scores better than a rushed complete one."

That last sentence protects your most conscientious candidates from working nine hours and resenting you.

What are planted issues, and why do they work so well?

They're deliberate flaws in the existing code or docs that the task touches. AI tools tend to build on top of what's already there. A candidate who reads the code and runs their work will notice. A candidate who accepts suggestions will extend the flaw.

Plant one of each kind:

Kind Example
Subtle bug A date filter that's off by one day because of a timezone conversion. A helper that swallows an error and returns an empty list
Misleading requirement A spec line that sounds clear until you look at the seed data. The right move is to question it and state an assumption
Security or performance trap An endpoint that doesn't check the record belongs to the user. A query that's fine with 10 rows and slow with 10,000

Two rules keep this fair rather than a gotcha:

  • Every issue must be discoverable from the repo and brief alone. If a strong engineer couldn't find it in the time box, it's a trick, not a test.
  • Nothing should crash immediately. That only tests whether they ran it once. You want to see whether they checked it.

And one scoring rule: flagging an issue in the README without fixing it earns most of the credit. Fixing it silently earns less, because you can't tell they noticed.

What's an AI usage log, and isn't it easy to fake?

It's a short table, 5–15 rows, written as they work. Not a transcript.

Tool What I asked for What I kept, changed or rejected How I checked it
Cursor Endpoint skeleton following the existing orders route Kept, renamed two fields to match conventions Ran the route tests
Claude Why the date test fails around midnight UTC Rejected its fix, which changed the test. Fixed the conversion instead Added tests at 23:30 and 00:30 UTC

Add three questions at the bottom: Where did AI save you the most time? Where was it wrong? Is there any part you don't fully understand?

Can it be faked? Sure. But the log isn't there to catch anyone. It's a map for the defense. It tells you which functions to ask about, and the "kept, changed or rejected" column is often the best evidence of judgment in the whole submission. A thin log is a reason to ask more in the defense, not a fail.

How do I run the oral defense?

Book 45–60 minutes. Read the code, README, log and walkthrough beforehand. Prepare two functions to ask about (at least one the log says the AI wrote), one requirement change, one perturbation, and a bug report for any planted issue they missed.

Tell candidates about the defense in the invite. It's fair warning, it costs honest candidates nothing, and it's the strongest anti-outsourcing measure you have.

A 45-minute version:

Minutes 0–5: set up.

"Please share your screen with your editor open. I'll ask you to walk me through parts of your code and then make a couple of small changes. Think out loud. Near the end I'll ask a few questions where you just talk, no tools. 'I don't know' is a perfectly good answer."

Minutes 5–12: explain.

"Take me through this function line by line. What does it do, and why is it written this way?"

Pick a function the log marks as AI-written. People who reviewed it explain the why. People who pasted it struggle.

Minutes 12–30: modify.

Change a requirement in a comment on screen:

"Product just told us discounts can't stack. One per order. Make that change."

Then a perturbation:

"Now the pricing service times out 5% of the time. What happens to an order, and what would you change?"

Watch how they move through their own code. Someone who wrote it, or seriously reviewed it, goes straight to the right file.

Minutes 30–40: planted issues.

If they found one: "How did you notice that?" If they missed one, hand over the bug report and a failing test, and watch them debug. Missing a planted issue in the take-home is common and not disqualifying. Debugging it live, methodically, is strong evidence.

Minutes 40–45: tradeoffs.

"What did you decide not to do? What would worry you about shipping this tomorrow?"

You can decide whether AI is allowed for the mechanical edits during the defense. Either way, the explaining is tools-off. That's the part that proves ownership.

How do I score it?

Score judgment, not output. Five criteria, each on a 1–4 scale with written anchors:

Criterion What a 4 looks like
Verification of AI output Log shows deliberate rejections. Explains AI-written code as fluently as their own
Detection of planted issues All three found, each fixed or flagged with reasoning, with a test for the bug
Tradeoffs made explicit Clear decisions with costs, and what they'd revisit
Tests Tests target the risky parts, including a regression test
Ownership in the defense Makes changes quickly, anticipates side effects, suggests improvements

Two scores should be disqualifying on their own: a 1 on verification (they shipped AI output unread) and a 1 on ownership (they can't change or explain what they submitted). In the second case you haven't seen their work, whatever the reason.

Don't score how much code the AI wrote, which tool they used, or feature count. A smaller, verified, well-tested solution beats a big unchecked one.

Should I really pay for a take-home?

Yes. Use a formula, not a negotiation: time box in hours times a fair hourly rate for the level, rounded up. A 3-hour task at $80 an hour is $240, so call it $250. Publish it in the invite, pay everyone the same, and pay within a week of submission, not on hire.

It widens your pool to strong engineers who already have jobs, it raises completion, and it lets you reasonably ask for tests, a README and a walkthrough. For candidates who can't accept payment because of employment or visa terms, offer a charity donation instead.

What about leaks and outsourcing?

Assume the task will leak eventually. Rotate the planted issues and one requirement every quarter or every ten candidates, and record which variant each person got.

Then add three cheap consistency checks: the person in the walkthrough is the person in the defense; the voice and level of the log, walkthrough and defense match; and the defense includes at least one change they couldn't have prepared for. None of these are verdicts. Inconsistency is a reason to add one more independent stage. If identity itself is the worry, see fake job candidates and proxy interviews.

Checklist: an AI-era take-home

  • [ ] AI explicitly allowed in the invite, with the policy linked
  • [ ] Scaffold repo, not a blank one, with tests that run in one command
  • [ ] Three planted issues, each discoverable from the repo alone
  • [ ] Time box set from an engineer's timed run, capped at 4 hours
  • [ ] Flat fee published and paid on submission
  • [ ] README, AI usage log and 5-minute walkthrough requested
  • [ ] Oral defense booked and announced in advance
  • [ ] Rubric scores verification, detection, tradeoffs, tests and ownership
  • [ ] Written walkthrough and extended time offered on request

What should I do next?

If you haven't written down your AI rules yet, start with the free AI interview policy template and the free cheat-risk audit. For the live rounds that sit alongside the take-home, read technical interview follow-up questions and should you allow AI in coding interviews.

If you'd rather not build scaffolds from scratch, Unscripted, the paid kit includes role-specific work samples (frontend, backend, full-stack, mobile, DevOps, data, ML) with planted issues, rubrics and defense scripts ready to adapt.

Hrishikesh Pardeshi is co-founder & CTO of Flexiple and has spent 10 years hiring and vetting developers. He wrote Unscripted, a vendor-neutral kit for running developer interviews in the AI era.