# AI purity test: the quiz and the real check

> What people mean by an AI purity test, why the viral score proves nothing, and how to check AI-generated code against behaviour your team agreed.

Source: https://cueapril.com/guides/ai-purity-test
Updated: 2026-10-07
Index: https://cueapril.com/llms.txt

## What is an AI purity test?

An AI purity test is a self-scored quiz that asks how much of your work you hand to AI, then gives you a number. It is entertainment: the answers are self-reported and nobody checks them. Cue April does not publish one. The testable version of the question is whether the AI-generated code in your repository behaves the way your team said it should.

Purity tests spread because they are quick, personal and slightly confessional. The AI version applies that format to how much you let a model do. It is a good prompt for a conversation, and it settles nothing about the software you ship. This guide covers both halves: what the quiz is, and what an honest check of AI-generated work looks like.

## What the viral quiz actually measures.

A purity test is a list you score yourself against: the more items you admit to, the lower your score. The AI version swaps the list for AI habits. Letting an agent write a function. Merging a diff you did not read. Asking a model to write the tests for the code it just wrote. That score is a talking point and nothing more. It is self-reported, unverified and weighted by whoever wrote the questions, so two people with the same habits can score differently on two different quizzes.

## The question underneath is worth testing.

Take the scoring away and a real worry remains: how much of this codebase do we still understand? That worry is testable, and the answer does not depend on who typed the code. Write down the behaviour the change should deliver, including its boundaries and the cases that should fail, then check the change against that record rather than against the tests that arrived with it. Cue April proposes a test plan for a code change, runs the checks your team approves across seven testing methods and returns each finding with the evidence behind it.

- Agree the expected behaviour before you read the generated tests.
- Run several methods, because each one reports only its own scope.
- Keep passed, failed, blocked and incomplete checks distinct.

## Use mutation testing as the honest score.

If you want a number, use one that something measured. Mutation testing makes small deliberate edits to your code and runs the tests against each one. A mutant that survives means no test noticed the change. The owned local sample published on this site records twelve generated mutants, six of them survivors, alongside 19 task executions, nine passed and ten failed. It is a local fixture with seeded faults and no cloud or model execution, so it is not a customer benchmark. It still makes the point: in that sample, half the deliberate code changes went unnoticed.

## No score decides your release.

A quiz score proves nothing, and a mutation score describes only the scope that ran. Passing checks describe the assertions that ran, not the behaviour nobody wrote a check for. Cue April does not decide releases: your team reviews the evidence and makes that call. An evaluation agrees the repository, the expected behaviour, the supported methods, the execution environment and the run allowance before anything executes.

## Questions

### Is there an official AI purity test?

No. Purity tests are an internet format rather than a standard, so anyone can publish one and choose its questions and scoring. There is no marking scheme, no governing body and no way to verify an answer. Treat any score you get as entertainment.

### Does an AI purity test detect AI-generated code?

No. It asks you what you did and inspects nothing. Authorship and correctness are separate questions too: even a confident guess about who wrote a change tells you nothing about whether it behaves the way your users need. Test the behaviour instead.

### What should we check instead of a score?

Check the change against the behaviour your team agreed: the user journey, the business rule, the API expectation and the cases that should fail. Then run the methods that suit it and read the surviving mutants. Cue April keeps the tested revision, the outcomes and the unfinished work in one report.

### Can you measure how much of our codebase AI wrote?

Cue April makes no claim about measuring authorship. It tests an approved change regardless of who or what wrote it. An authorship figure belongs to your commit history and review process, not to a test run.

## Related

- [Testing AI-generated code](https://cueapril.com/guides/testing-ai-generated-code): How to test AI-generated code: check it against requirements your team agreed, with several test methods, mutation checks and inspectable evidence.
- [Mutation testing](https://cueapril.com/testing/mutation-testing): Mutation testing shows which deliberate code changes your tests miss. Run Stryker mutation testing and inspect the native results and survivors.
- [Agentic testing](https://cueapril.com/guides/agentic-software-testing): Agentic testing explained: approved intent, coordinated test methods, reproducible findings and evidence for a human release decision.
- [Pull request testing](https://cueapril.com/use-cases/pull-request-testing): Connect approved testing to GitHub pull requests. Review tested revisions, actual outcomes, findings and outstanding scope before deciding to merge.
