# Auditing a codebase for unlicensed copied code

This is a working note on auditing a codebase for unlicensed copied code: what held up over several terms, what quietly stopped being done, and which of those two was actually a problem.

## Deciding what the output is for

Distinguish between code that is similar and code that shares a history. Two implementations of the same textbook algorithm are similar and unrelated. Two files with the same unusual variable ordering, the same dead branch and the same off-by-one comment share a history. Systems that report only the former make the reviewer do the work of finding the latter, on every single pair, forever.

## What the reviewer actually needs

Publish the method to the people being measured. Students and candidates who know what is compared, against what, and what happens next behave differently from those who do not — and the difference shows up as less of the thing you were detecting. Detection and deterrence are not in tension here; secrecy about the method buys a marginally higher catch rate and gives up nearly all of the deterrent effect.

## Making it survive the year

Version the corpus, not just the code. A comparison run in March against a corpus that has since grown cannot be reproduced in June, and "we re-ran it and got a different number" is a sentence that ends processes. Recording which snapshot a result came from costs a column in a table, and is the difference between a finding that survives review and one that evaporates under it.

## Putting it into practice

What separates a [source code plagiarism detection](https://codequiry.com) from a diff is that it can tell you what the overlap means.
