Most AI detectors use generative AI to catch generative AI.
That's circular — and it's why they fail. Proofline measures structure instead: no generative model, no reasoning, anywhere in the detection path.
Out of 100 samples of AI generated text, it catches 99.9 of them.
*7,400 sample poolOut of 100 samples of human written text, it clears 95 and wrongly flags 5.
Less than 2 in 100 false positives when used together.
*tested on 1,800 documentsAsk a language model whether a language model wrote something, and you've built a detector that inherits every one of that model's blind spots — including a documented bias against non-native English writers, and drift every time the underlying model updates.
Read the full argument →No generative model in the measurement path.
Proofline computes structure directly from the characters — arithmetic, not inference. No inference, no learning model for meaning, no generative technology. Same text, same numbers, any machine, forever.
Signed, verifiable results
Every result is Ed25519-signed and reproducible on your own machine — not just downloadable.
Verify it yourself →No AI, no GPU, no data center
Runs on any CPU. Detecting AI text shouldn't require the same infrastructure as generating it.
See the case →Full methodology
The benchmark, the reference set, the per-configuration breakdown — nothing summarized away.
Read the proof →Built for the person writing, and the institution asking.
For writers & researchers
Check your own work before you submit it anywhere. Free to start, then $24.99/mo.
Get started free →For universities & publishers
Deployed on-prem or in your VPC. We're establishing the protocols to reach formal compliance certification — documentation on request. $249.99 to get started.
Talk to us →Try it on your own text
3 free scans a day. No account needed to start.
Pasted text (800 characters or more) is English-only for now; document upload scores any language.
Two documented failures in the detectors you're using now
Non-native English writers flagged as AI far more often
A fairness problem documented across detectors that score by fluency rather than structure.
Most leaderboard detectors never report non-English results
The public benchmarks buyers cite often skip the languages most likely to trigger false positives.