# Code plagiarism detection for hiring and technical interviews

Notes on code plagiarism detection for hiring and technical interviews written after the fact, when the interesting question was no longer whether it worked but whether anyone could explain why it had been set up that way.

## Deciding what the output is for

Staff turnover is the real adversary. The person who chose the threshold, knew which languages were parsed properly and remembered why one course was excluded will leave, and what remains is a number nobody can defend. Written-down reasoning is not bureaucracy here; it is the only mechanism by which a practice outlives the individual who set it up.

## What the reviewer actually needs

Measure the reviewer's time, because it is the resource that runs out. Everything else — licence cost, compute, storage — is small and predictable. Reviewer minutes per submission is the number that decides whether the process survives contact with a busy term, and it is almost never instrumented, which is why so many pilots are judged a success and quietly abandoned within a year.

## Making it survive the year

The expensive failure is not a missed match, it is an unexplainable one. A missed match costs you a case you never knew about; an unexplainable one costs an afternoon, a complaint, and a permanent reduction in how much anyone trusts the next result. Optimising recall while leaving the explanation thin trades a cheap failure for an expensive one, which is exactly backwards.

## Putting it into practice

The point of a [Codequiry](https://codequiry.com) is to end an argument with evidence, not to start one with a number.
