What is an AI purity test?
An AI purity test is a self-scored quiz that asks how much of your work you hand to AI, then gives you a number. It is entertainment: the answers are self-reported and nobody checks them. Cue April does not publish one. The testable version of the question is whether the AI-generated code in your repository behaves the way your team said it should.
What the viral quiz actually measures.
A purity test is a list you score yourself against: the more items you admit to, the lower your score. The AI version swaps the list for AI habits. Letting an agent write a function. Merging a diff you did not read. Asking a model to write the tests for the code it just wrote. That score is a talking point and nothing more. It is self-reported, unverified and weighted by whoever wrote the questions, so two people with the same habits can score differently on two different quizzes.
The question underneath is worth testing.
Take the scoring away and a real worry remains: how much of this codebase do we still understand? That worry is testable, and the answer does not depend on who typed the code. Write down the behaviour the change should deliver, including its boundaries and the cases that should fail, then check the change against that record rather than against the tests that arrived with it. Cue April proposes a test plan for a code change, runs the checks your team approves across seven testing methods and returns each finding with the evidence behind it.
- Agree the expected behaviour before you read the generated tests.
- Run several methods, because each one reports only its own scope.
- Keep passed, failed, blocked and incomplete checks distinct.
Use mutation testing as the honest score.
If you want a number, use one that something measured. Mutation testing makes small deliberate edits to your code and runs the tests against each one. A mutant that survives means no test noticed the change. The owned local sample published on this site records twelve generated mutants, six of them survivors, alongside 19 task executions, nine passed and ten failed. It is a local fixture with seeded faults and no cloud or model execution, so it is not a customer benchmark. It still makes the point: in that sample, half the deliberate code changes went unnoticed.
No score decides your release.
A quiz score proves nothing, and a mutation score describes only the scope that ran. Passing checks describe the assertions that ran, not the behaviour nobody wrote a check for. Cue April does not decide releases: your team reviews the evidence and makes that call. An evaluation agrees the repository, the expected behaviour, the supported methods, the execution environment and the run allowance before anything executes.