An AI score is a classifier's reading of how predictable your writing is, printed as a number. It is not the probability that you used AI, and 100% does not mean every word came from a machine. A detector never sees how a document was written. It sees finished text and measures how closely each word matches what a language model would have put in that position. A high number means the writing is unsurprising. Unsurprising writing comes out of machines, and it also comes out of careful editing, second-language instruction, formal registers, and anyone who followed the rubric closely. That is why the same paragraph can read as machine-like on one tool and human on another, and why a score is a reason to look closer rather than a finding on its own.
Detectors are statistical classifiers. They estimate, word by word, how likely each word was given the words before it, then aggregate that into one figure. Text the model finds easy to predict reads as machine-written. Text that surprises it reads as human. That is the entire mechanism. The classifier reads finished prose and nothing else. It has no record of which keys you pressed and no copy of the prompt you may or may not have used.
The figure you see sits at the end of a chain of conversions. The classifier produces a raw measurement on its own internal scale. The vendor chooses the point at which that measurement counts as a flag. Then it gets mapped onto something readable, which is where percentages and coloured bars come from. Vendors label the result differently: some present a confidence, some the share of sentences or passages that crossed a threshold, some only a band. All three look identical on screen and mean different things, so read the label the vendor puts on its own number before you read the number.
None of this is authorship evidence. A detector has no source document, no drafting history, and no chain of custody. It has your prose and a model of ordinary language.
Two questions get collapsed into one. The detector answers: how machine-like is this writing? The reader hears: how likely is it that this person cheated? Moving from the first answer to the second requires two things the score does not contain, namely how often human writing looks like this, and how much of the pile was human to begin with.
Take a stack of essays in which most were written by hand. Even if a detector catches nearly every machine essay and misfires on only a small share of the human ones, the flagged pile still holds human essays, because the human pile was so much larger at the start. That is the shape of the arithmetic. When almost everything you scan is human, a confident-looking flag carries less weight than it appears to.
The misfires are also not spread evenly. Liang et al. at Stanford HAI in 2023 found that 61% of TOEFL essays written by non-native speakers were falsely flagged as AI across seven detectors. Plain, correct, evenly paced English is the most predictable English there is, which is precisely what the measurement rewards.
One more thing 100% does not mean: on tools that report the share of flagged segments, a whole-document 100 can come from a short document in which every segment happened to clear the line. It is a statement about segments, not about how much of your thinking was yours.
Nothing standardises these numbers. There is no shared scale, no shared test set, and no shared definition of what counts as a flag. Two detectors run on the same paragraph on the same afternoon are two different models trained on different text, compared against cut-offs that two different companies chose privately. Disagreement is the expected result, not a malfunction in one of them.
The differences that produce the spread:
The checker here returns two things: a number from 0 to 100, where higher means the text reads more machine-like, and one of three verdicts. Those verdicts are "Reads as AI-generated", "Borderline", and "Reads as human-written". The verdict is the reading. The number tells you where you sit inside it and which way an edit moved you.
The score ranks how machine-like the text reads. It is not a probability, which is why it is called a score and never a percentage chance. Reading it as "42% of my essay is AI" or "a 42% chance I get accused" imports a meaning it does not carry. Read the direction and the size of the move rather than the digits. A couple of points either way is noise.
Borderline means unresolved. It sits between the two readings, in the range where the evidence does not settle the question. Treat it as neither result rather than as a pass.
Two practical notes. The check needs at least 65 words, because a shorter passage gives the measurement too little to read. And a check costs 1 credit whatever the length, against 10 free credits a day per network, with no account, no login, no card and no email address to verify. Checking is priced flat so that measuring twice, once before an edit and once after, is the normal way to use it.
Because the measurement is predictability, the features that move it are the phrasings a model reaches for by default. The AI vocabulary checker runs in your browser, is free, has no run limit, and highlights the exact words rather than handing you a number. Run it alongside the check and you get the diagnosis and the location together.
The edits that reliably shift the reading:
If you would rather have the rewrite done for you, the humanizer rewrites every sentence in one click, up to 12,000 characters a run, at 1 credit per 50 words. On policy: rewriting a draft you were permitted to generate is ordinary editing, and where AI is not permitted, that rule covers the draft however it is edited. Read the policy that applies to you.
Rewriting your own honest draft to satisfy a classifier is not the only move available to you. Check the draft, read the highlighted tells, and see which of your own habits are being scored. In most flagged essays it turns out to be two or three repeated constructions running through the whole piece rather than anything wrong with the argument.
If someone has put a number in front of you, two questions are worth asking before you answer the accusation. What does the vendor say this number is, a confidence or a share of flagged passages? And what threshold turned it red, chosen by whom? A percentage without its cut-off is a colour, not a finding.
The stronger answer is evidence rather than argument. The signed drafting record here keeps a signed history of a document as you write it, using ECDSA P-256, and a third party can verify that signature without contacting us. It works on the essay you are about to write, which is the reason to start it now rather than after the meeting.
It means the text sat at the far end of that particular tool's scale, at the point the vendor labels as fully machine-like. It does not mean 100% of your document was generated, and it is not a 100% probability that you used AI. On tools that report the share of flagged passages, a document-level 100 only means every passage crossed a line the vendor chose. The reading is about how predictable your phrasing is, and nothing in it observed how the document was written.
No. A detector has no source document, no drafting history and no chain of custody. It measures writing style and infers from it. A high score is a reasonable prompt to look more closely at a piece of work. It cannot stand in for what you find when you do, and it cannot separate machine text from a careful writer whose prose happens to be predictable.
Because there is no shared scale. Each tool uses a different base model trained on different text, picks its own threshold for a flag, splits the document differently, and maps the result onto 0-100 its own way. Vendors also update their models, so the same text in the same tool can read differently later. Disagreement between tools is the normal outcome rather than a sign that one of them broke.
Aim for the verdict rather than a target figure. Here that means moving out of "Reads as AI-generated" and past "Borderline", which is the band where the answer is unresolved and should not be read as a pass. Every tool sets its own scale and its own cut-off, so a number chased on one scale does not transfer to another. Fix the phrasing that made the text predictable and the reading follows.
No, and often the opposite. The measurement rewards surprise, so clean, correct, evenly paced prose scores as machine-like. Second-language writers, technical writers and students who followed a rubric closely all produce exactly that. The Stanford HAI study by Liang et al. in 2023 found 61% of TOEFL essays by non-native speakers falsely flagged across seven detectors. What the score tells you is that your phrasing is predictable, which is a different observation from whether the writing is any good.
Sixty-five words minimum, because shorter passages do not give the measurement enough to read. Each run costs 1 credit whatever the length, and the free tier is 10 credits a day per network with no account, no login, no card and no email address. The proofreader, word counter, readability checker, AI vocabulary checker, originality check and signed drafting record are free and unlimited, and run in your browser.