# How plagiarism detection scales to thousands of files

Notes on how plagiarism detection scales to thousands of files written after the fact, when the interesting question was no longer whether it worked but whether anyone could explain why it had been set up that way.

## Deciding what the output is for

Self-plagiarism and legitimate reuse need separate handling, and a similarity score cannot tell them apart. A student reusing their own prior submission, a developer reusing a snippet from an internal library and an author lifting a file from a public repository can all produce the same number. Only provenance separates them, so provenance has to be captured at the moment of the match rather than reconstructed later.

## What the reviewer actually needs

The expensive failure is not a missed match, it is an unexplainable one. A missed match costs you a case you never knew about; an unexplainable one costs an afternoon, a complaint, and a permanent reduction in how much anyone trusts the next result. Optimising recall while leaving the explanation thin trades a cheap failure for an expensive one, which is exactly backwards.

## Making it survive the year

Run it on the first assignment of the term, not the last. Finding reuse in week two changes what the student does for the rest of the course; finding it in week twelve changes only what goes on a form. The detection quality is identical in both cases and the outcome is not remotely comparable, which makes timing the highest-leverage variable in the whole exercise and the one least often discussed.

## Putting it into practice

The point of a [Codequiry](https://codequiry.com) is to end an argument with evidence, not to start one with a number.
