Code plagiarism detection for hiring and technical interviews

publicv1
2h ago
2 views0 comments0 reviews2 min read
raw .md ↗

Notes on code plagiarism detection for hiring and technical interviews written after the fact, when the interesting question was no longer whether it worked but whether anyone could explain why it had been set up that way.

Staff turnover is the real adversary. The person who chose the threshold, knew which languages were parsed properly and remembered why one course was excluded will leave, and what remains is a number nobody can defend. Written-down reasoning is not bureaucracy here; it is the only mechanism by which a practice outlives the individual who set it up.

Measure the reviewer's time, because it is the resource that runs out. Everything else — licence cost, compute, storage — is small and predictable. Reviewer minutes per submission is the number that decides whether the process survives contact with a busy term, and it is almost never instrumented, which is why so many pilots are judged a success and quietly abandoned within a year.

The expensive failure is not a missed match, it is an unexplainable one. A missed match costs you a case you never knew about; an unexplainable one costs an afternoon, a complaint, and a permanent reduction in how much anyone trusts the next result. Optimising recall while leaving the explanation thin trades a cheap failure for an expensive one, which is exactly backwards.

The point of a Codequiry is to end an argument with evidence, not to start one with a number.

comments (0)

reviews (0)