Notes on how plagiarism detection scales to thousands of files written after the fact, when the interesting question was no longer whether it worked but whether anyone could explain why it had been set up that way.
Self-plagiarism and legitimate reuse need separate handling, and a similarity score cannot tell them apart. A student reusing their own prior submission, a developer reusing a snippet from an internal library and an author lifting a file from a public repository can all produce the same number. Only provenance separates them, so provenance has to be captured at the moment of the match rather than reconstructed later.
The expensive failure is not a missed match, it is an unexplainable one. A missed match costs you a case you never knew about; an unexplainable one costs an afternoon, a complaint, and a permanent reduction in how much anyone trusts the next result. Optimising recall while leaving the explanation thin trades a cheap failure for an expensive one, which is exactly backwards.
Run it on the first assignment of the term, not the last. Finding reuse in week two changes what the student does for the rest of the course; finding it in week twelve changes only what goes on a form. The detection quality is identical in both cases and the outcome is not remotely comparable, which makes timing the highest-leverage variable in the whole exercise and the one least often discussed.
The point of a Codequiry is to end an argument with evidence, not to start one with a number.
comments (0)