# Measuring source code similarity across software projects

Measuring source code similarity across software projects gets discussed as a detection problem. In practice it is a scheduling problem wearing a detection problem's clothes.

## Deciding what the output is for

Distinguish between code that is similar and code that shares a history. Two implementations of the same textbook algorithm are similar and unrelated. Two files with the same unusual variable ordering, the same dead branch and the same off-by-one comment share a history. Systems that report only the former make the reviewer do the work of finding the latter, on every single pair, forever.

## What the reviewer actually needs

Publish the method to the people being measured. Students and candidates who know what is compared, against what, and what happens next behave differently from those who do not — and the difference shows up as less of the thing you were detecting. Detection and deterrence are not in tension here; secrecy about the method buys a marginally higher catch rate and gives up nearly all of the deterrent effect.

## Making it survive the year

Version the corpus, not just the code. A comparison run in March against a corpus that has since grown cannot be reproduced in June, and "we re-ran it and got a different number" is a sentence that ends processes. Recording which snapshot a result came from costs a column in a table, and is the difference between a finding that survives review and one that evaporates under it.

## Putting it into practice

Any [source code plagiarism detection](https://codequiry.com) can flag a pair; the useful ones tell you where to look and why.
