The edits that move an AI detection score are structural rather than lexical. A statistical detector is not reading your text for meaning, and it is not matching your words against a banned list. It scores how predictable the writing is, word by word, and how much that predictability rises and falls from one sentence to the next. The ranking of tactics falls out of that. Swapping single words for synonyms does almost nothing, because the shape of the sentence survives the swap. Planting deliberate errors may shift the number slightly, and it costs you marks with the person grading the work. What changes a reading is varying sentence length, cutting stock openers and transitions that connect nothing, and breaking the even paragraph shape a model produces by default. One point of policy before the technique. Where AI is not permitted, that rule covers the draft however it is edited, so read the policy that applies to you.
Two statistics sit under the common approach to AI classification. The first is average predictability. Given the words so far, how expected was the next one? A language model is built to pick likely continuations, so its output runs smooth. The second is how much that predictability varies across the passage. Human drafts spike and dip. A writer reaches for an odd word, cuts a sentence short, spends a clause on something that turns out not to matter, then gets back on track.
Both numbers describe the whole passage rather than any phrase inside it. That is why you can strike every word on a list of AI vocabulary and watch the score barely move. The sentence lengths did not change. The clause order did not change. The rhythm did not change. Structure is the thing to edit, because structure is what those numbers are computed from.
The same arithmetic explains why a high score is not evidence about who wrote something. Careful, even, grammatically clean prose can read as machine-made whether a machine made it or not. Liang and colleagues at Stanford HAI reported in 2023 that 61% of TOEFL essays by non-native English speakers were falsely flagged as AI across seven detectors. If you wrote it yourself and got flagged anyway, the false-positives page is the better use of your time.
Most of the popular advice operates below the level the score is computed at. Each tactic edits words while leaving the structure that produced the reading intact. None of them is worth much of your time.
The most repeated tip is to add typos, splice a few commas, drop an article here and there. It can shift the reading, for a reason that should stop you. It raises average surprise by making the writing worse. Every error you plant is a mark you give up with the human being who reads the work and decides what it is worth. You are trading a number you cannot see for a grade you can.
There is a second cost. Slightly-off English is the signature of a second-language writer, and that is the group the Stanford figures show being falsely flagged. Broken prose does not read as human. It reads as careless.
A workable rule: do not make an edit you would not defend out loud to the person grading the work. Everything below passes that test, because a good editor would ask for it anyway.
Model output clusters. Left alone it produces sentence after sentence of similar length, each one built as subject, verb, object, qualifier. Human writing does not hold a band. A paragraph might run four words, then thirty-one, then nine. Count the words in each sentence of a paragraph you are worried about. If every count lands within a few of every other, fix that before you touch a single word choice.
Three edits do most of the job. Split a long sentence at its conjunction and let one half stand alone as a short one. Fold two related short sentences into a single longer one with a subordinate clause. Delete the sentence that restates the sentence before it, which a model adds by habit and which you will find in almost every paragraph it writes.
Here is the shape of it. Say three sentences run 21, 24 and 20 words. After those edits they run 6, 33 and 11. No error was introduced and no claim changed. The paragraph now has a pulse. Read it aloud and you will hear the difference before any tool tells you, which is the cheapest way to find the flat stretches in a long document.
The second lever is deletion. Models fill slots, and some of those slots produce phrases that carry no information. Removing them cuts the runs of highly predictable words and tightens the prose at the same time. The version without them is better writing.
The registry on this site names these tells and highlights the exact phrases where they occur in your own text, free and unlimited. The categories worth knowing by heart are short.
Above the sentence there is a third pattern most people never touch. Model paragraphs come out at roughly equal length, each with a topic sentence, two supporting sentences and a closer. The sections come out equal too. Real writing is lumpy because evidence is lumpy. The point with three sources gets a long paragraph, and the point that needs one line gets one line. Put the headings where the argument turns rather than at even intervals.
Then measure, because guessing is what makes this slow. Paste 65 words or more into the checker on this page and you get a score from 0 to 100 in seconds, along with one of three readings: reads as AI-generated, borderline, or reads as human-written. Change one thing, run it again, and you learn which of your own habits is carrying the weight. That feedback loop is the real skill, and it transfers to everything you write next.
If you would rather start from a rewrite, the humanizer restructures every sentence in one click, up to 12,000 characters per run, which is roughly 2,000 words. Longer documents go through in consecutive runs. The check costs one credit at any length, the rewrite costs one credit for every 50 words, and the free tier is 10 credits a day with no account, no card and no email address. Reading the tells, the proofreader, the word counter and the readability check stay free and unlimited.
Nobody can promise you a particular result from a particular detector, and a tool that promises one is selling something it cannot deliver. What you can do is change the properties detectors measure, which are predictability and its variance. Sentence-length variance, deleted stock phrasing and uneven paragraph structure move those properties. Word swaps and typos mostly do not.
A paraphraser that keeps the sentence structure and trades words for synonyms changes very little, because the reading is computed over sequences and rhythm rather than over individual word choices. It also tends to make the prose slightly worse, since the second-choice synonym is usually the less precise one. Changing the structure does more than changing the surface.
No. The effect is small and the cost is not. You are handing away marks with the human reader in order to move a number you cannot verify, and you are making your writing look like the second-language prose that the Stanford study found being falsely flagged. Fix the structure and leave the grammar alone.
Different detectors are built on different training data and weigh different features, so their numbers do not line up and no single score is a verdict. What holds across them is the underlying property. Text that is uniformly predictable reads as machine-made, and text with real variation in sentence length and rhythm does not. That is the property worth working on, since it is the one every classifier of this kind is looking at in some form.
That is a different problem with its own answer. Detectors flag clean, even, correct prose, which is exactly what careful writers and second-language writers produce. The Stanford HAI study found 61% of TOEFL essays by non-native speakers falsely flagged across seven detectors. Start with the false-positives page and the page for people who have been accused, which cover what evidence you can assemble and how to talk about it.
The check needs at least 65 words and costs one credit per run at any length. The rewrite handles 12,000 characters per run, about 2,000 words, from a 25-word minimum, and costs one credit for every 50 words. The free tier is 10 credits a day per network with no account and no card. The tell registry, proofreader, word counter, readability check, AI vocabulary check and originality check are free and unlimited.