Notes on how to detect source code plagiarism in programming assignments written after the fact, when the interesting question was no longer whether it worked but whether anyone could explain why it had been set up that way.
Decide what the output is for before choosing what produces it. A number that feeds a conversation, a report that feeds a formal process and a log that feeds an audit have almost nothing in common, and a system tuned for one is actively unhelpful for the others. Most disappointment traces back to a tool bought for the third purpose being used for the first, by people who were never told which it was.
Run it on the first assignment of the term, not the last. Finding reuse in week two changes what the student does for the rest of the course; finding it in week twelve changes only what goes on a form. The detection quality is identical in both cases and the outcome is not remotely comparable, which makes timing the highest-leverage variable in the whole exercise and the one least often discussed.
Measure the reviewer's time, because it is the resource that runs out. Everything else — licence cost, compute, storage — is small and predictable. Reviewer minutes per submission is the number that decides whether the process survives contact with a busy term, and it is almost never instrumented, which is why so many pilots are judged a success and quietly abandoned within a year.
The point of a Codequiry is to end an argument with evidence, not to start one with a number.
comments (0)