Skip to content
← All notes

Can you still trust a take-home assignment in the AI era?

Not as a test of whether somebody can produce the artefact, because everybody can now. It still works, but only if you change what the exercise is measuring.

Hiring7 min read

The standard take-home was built on an assumption that no longer holds: that producing a good solution required being able to produce a good solution.

Every candidate can now submit competent work. So the exercise no longer separates anybody, and the usual reactions are to abandon take-homes entirely or to try to detect AI use. Both are worse than the alternative.

Why detecting AI use is the wrong goal

Detection does not work reliably, it produces false accusations that damage good candidates, and it disproportionately catches non-native speakers whose writing already looks unusual to a detector.

It is also aiming at the wrong target. Using AI is not cheating at a job where using AI is part of the job. What you actually want to know is whether they understand what they submitted, and that is answerable without any detection at all.

Detection aimed at the wrong question. Whether a tool was used is not the thing worth knowing; whether the person understands what they handed in is, and that can be established directly.
Using the tools is not cheating at a job that uses the tools.

Redesign the exercise around judgement, not production

The fix is to make the interesting part of the task something a model cannot do for the candidate, and then to ask about it.

  • Include a deliberate ambiguity. The good candidate names it and says how they resolved it. A generated answer picks one silently and does not mention that it chose.
  • Include a constraint that makes the obvious answer wrong. Models produce the common solution, which is exactly the failure you want to see.
  • Ask for what they rejected and why. Understanding lives in the discarded options and generated work rarely has any.
  • Ask what they would do differently with two more weeks. This separates people who understand the tradeoffs from people who produced an artefact.

Always follow it with a conversation

This is the part that does the real work, and it takes twenty minutes. Walk through their submission and ask why in a few places. Then change one requirement and ask what would have to change in response.

Somebody who understands their own submission handles this comfortably. Somebody who does not runs out of answers almost immediately, and it is not subtle. No detection tool is needed and no accusation has to be made.

A twenty minute conversation about a submission, changing one requirement and asking what that changes. Understanding survives the follow-up questions and its absence does not.
The artefact is the shared object. The conversation is the assessment.

And be honest about the cost you are imposing

Take-homes were always a tax on candidates with the least free time, which usually means people with caring responsibilities or a second job. That was true before and it is still true.

If the artefact no longer distinguishes anybody, a long take-home is a large cost for very little information. Keep it short, be explicit that AI use is fine, and put the real assessment in the conversation, which is faster and tells you more.

Common questions

Are take-home assignments still worth using?
Yes, if you change what they measure. As a test of whether somebody can produce the artefact they are worthless, because everybody can. As a shared object to have a detailed conversation about, they are still one of the better tools available.
Should candidates be allowed to use AI on a take-home?
Say yes explicitly, because they will anyway and the ones who follow the rules are the ones you penalise by pretending otherwise. Then design the exercise so understanding is what gets tested, and assess it in a follow-up conversation.
How do you tell if a candidate actually understands their submission?
Change something and ask what that changes. Ask what they considered and rejected. Ask what the weakest part is. Somebody who did the thinking answers comfortably; somebody who did not runs out of answers within a couple of follow-ups.
Do AI detection tools work for hiring?
Not reliably enough to make decisions on. They produce false positives, and they hit non-native speakers hardest, which turns a detection error into a discrimination problem. Assessing understanding in conversation answers the same question without the risk.

PDP Quest exists because of the problem underneath all of these: when output stops indicating capability, you need another way to know who can actually do the work.

See how verification works →