# How similarity engines normalize code

A short account of how similarity engines normalize code from the side that has to sit in the meeting afterwards rather than the side that runs the comparison.

## Deciding what the output is for

Take the false positives seriously as a design input. Every pattern that reliably produces a harmless high score — generated code, a shared template, a language whose idioms are narrow — is something the pipeline can be told about once. Teams that log why each dismissal happened end the year with a filter that makes the next year's queue a third shorter. Teams that dismiss and move on start every year from the same place.

## What the reviewer actually needs

Decide what the output is for before choosing what produces it. A number that feeds a conversation, a report that feeds a formal process and a log that feeds an audit have almost nothing in common, and a system tuned for one is actively unhelpful for the others. Most disappointment traces back to a tool bought for the third purpose being used for the first, by people who were never told which it was.

## Making it survive the year

Staff turnover is the real adversary. The person who chose the threshold, knew which languages were parsed properly and remembered why one course was excluded will leave, and what remains is a number nobody can defend. Written-down reasoning is not bureaucracy here; it is the only mechanism by which a practice outlives the individual who set it up.

## Putting it into practice

Corpus, provenance and a readable report — a [code similarity checker](https://codequiry.com) without all three is a demo.
