Justify

Does the AI assistant actually pay off?

Every line has to earn its own place.

Justify reads a whole repository and asks every import, function, class and dependency one question: why are you here? What cannot answer is dead weight. Justify finds it, shows how much came from AI-assisted commits, and removes it only with proof.

Try

  • Public Python repositories
  • Read-only — never runs your code
  • Open source, MIT

What it finds

Four kinds of weight. Every one passes every test.

Dead weight never breaks anything — that is why it survives. An assistant writes for the prompt in front of it, not for the codebase it cannot see, so it leaves these behind:

Unused imports

A module loaded that nothing calls. Every one is read, parsed and loaded on every run.

import io          # nothing uses it

Unused functions

A helper with no callers anywhere in the repository — kept alive only because nobody checked.

def clock_time():   # 0 callers

Duplicate helpers

The same job written twice, because the assistant could not see the first one.

def fmt_date(d) ≡ def format_date(d)

Unused dependencies

A package installed that no file imports — one more thing to download, patch and audit.

leftpad==1.0     # no file imports it

How it works

Seven stages. The model never deletes.

A code graph proposes, a model is asked why, a second call tries to prove it wrong, the tests decide, and a person approves. Doubt always means keep.

  1. 1IngestEvery file read once and fingerprinted.
  2. 2Static factsSyntax trees, not text search: who uses what, across the whole repository.
  3. 3CandidatesFrameworks, dynamic lookups and re-exports are held back for judgement.
  4. 4JustifyA model is asked why the unit exists, and must cite a file and line that is then checked.
  5. 5ChallengeA second call tries to prove the code is needed. If it can, the code stays.
  6. 6ProofRemoved in a temporary copy; your tests run. A file the tests never load is “not provable”.
  7. 7ReportA reason on every line, a ledger of every run. A person approves.

This site stages 1–3, authorship and the score — public repositories, read-only.

Your machine stages 4–6 — your model, your tests, your private code.

What we measured

Does it pay off? It depends on the repository.

We ran Justify on five public repositories with AI-assisted commits — 3,247 files, 989,614 lines — and split every finding by who wrote it.

  1. Unused code is not where AI-assisted code costs. It carried no more than human code in all five.
  2. Duplication is. Higher in three of five — 7.7× the human rate in the largest repository.
  3. There is no single answer. The same assistant pays off in one repository and costs in another. Measure yours.

Static stages and authorship, no proof. AI-assisted = commits with an assistant trailer, a lower bound. Rework is counted from the first AI-assisted commit.

Duplicate code per 1,000 lines
AI-assistedHuman
Duplicate code per 1,000 lines, AI-assisted versus human

Model Context Protocol

Let your AI assistant check its own work.

The assistant that writes the code can call Justify before it hands the code over. One URL — no account, no key.

Private code, and proof

The hosted audit reads public repositories and never runs anything. To audit private code, and to prove each removal by running your own tests on a temporary copy, run Justify on your machine:

pip install "justify-code[mcp] @ git+https://github.com/BPSKartik/justify"
justify scan . --prove "python -m pytest -q"

As a local MCP server: command justify-mcp, with JUSTIFY_TEST_COMMAND (the only command it may run) and JUSTIFY_ALLOWED_ROOTS (the only folders it may read).

Questions

Before you trust it.

Does Justify ever delete my code?

No. It proposes. On your machine it removes code only in a temporary copy, runs your tests there, and writes a report. A person approves every change. The model can veto a removal; it can never force one.

What does this website do with a repository?

It clones the public repository, reads it, traces each line to an AI-assisted or human commit, and deletes the clone. It never runs the repository's code or tests and never sends it to a model. Results are cached by commit.

How does it know which code an AI wrote?

From commit trailers the assistants write themselves, like Co-Authored-By: Claude or Co-authored-by: Copilot. An assistant used without a trailer counts as human, so the AI share is always a lower bound.

What is the Justified Line Ratio?

The share of lines that are not dead weight: lines left after removing everything with no use anywhere, divided by all lines. A team can track it per repository, per quarter, and split by authorship.

How is this different from a linter?

A linter looks at one file and asks whether a name is used. Justify looks at the whole repository, asks whether the code is worth keeping, explains why with a file and line, proves removals with your tests, and splits the result by who wrote it.

Which languages?

Python today. The parsing layer is built to take tree-sitter for more languages next.