How to bypass AI detection: what actually changes the reading

Most advice on bypassing AI detection does not work, and the advice that does work is the same on every guide that actually tests its methods: change the rhythm and the shape of the prose, not the words. Sentence-length variation, a broken-up paragraph structure, and specific detail carry a score, while synonym swaps and planted typos barely move it. The tools that promise a guaranteed human reading are promising something about a detector they do not run, so treat that claim as advertising. The honest version of this page is that no tool, including the one on this site, can promise you a particular score on a particular detector. What you can do is measure where a draft stands, change the things the measurement reads, and watch the number move. Paste your text into the checker on this page, press check, and you get a score from 0 to 100 in seconds, with one of three verdicts: reads as AI-generated, borderline, or reads as human-written. Then edit the things this page describes and check again. The check costs one credit at any length, the humanizer costs one credit per 50 words, and the free tier is 10 credits a day with no account, no card and no email address. One point of policy before the technique: where AI is not permitted, that rule covers the draft however it is edited, so read the policy that applies to you.

Try an example:
113 words
This run costs 3 · 10 of 10 free credits left today

The advice that survives testing

Every guide that actually tests its methods ends up recommending the same few edits, which is worth taking seriously because the guides disagree about almost everything else. The consistent list is short: vary sentence length, break the even paragraph shape, cut the stock phrases, add specifics that only you would know. That is the whole of it, and the same guides agree on what does not work, which is a longer list.

The rest of the advice circulating online is thinner. Some of it is harmless but useless, some of it makes the writing worse, and a surprising amount of it is advertising for a tool that promises a guaranteed result. A tool that promises a particular score on a detector it does not run cannot keep the promise, so treat the guarantee as marketing and keep the parts that survive testing: the structural edits below, in the order that moves the number most.

What a detector is actually measuring

Two statistics sit under the common approach to AI classification. The first is average predictability. Given the words so far, how expected was the next one? A language model is built to pick likely continuations, so its output runs smooth. The second is how much that predictability varies across the passage. Human drafts spike and dip. A writer reaches for an odd word, cuts a sentence short, spends a clause on something that turns out not to matter, then gets back on track.

Both numbers describe the whole passage rather than any phrase inside it. That is why you can strike every word on a list of AI vocabulary and watch the score barely move. The sentence lengths did not change. The clause order did not change. The rhythm did not change. Structure is the thing to edit, because structure is what those numbers are computed from.

The same arithmetic explains why a high score is not evidence about who wrote something. Careful, even, grammatically clean prose can read as machine-made whether a machine made it or not. Liang and colleagues at Stanford HAI reported in 2023 that 61% of TOEFL essays by non-native English speakers were falsely flagged as AI across seven detectors. If you wrote it yourself and got flagged anyway, the wrongly accused page is the better use of your time.

The tactics that barely move the number

Most of the popular advice operates below the level the score is computed at. Each tactic edits words while leaving the structure that produced the reading intact. None of them is worth much of your time.

Notice what the list has in common. Every item changes words, formatting or luck while leaving the sentence order, the clause shapes and the rhythm untouched, and those are the properties the reading is computed from. The wilder suggestions, like sneaking lookalike letters into the text, fail the simplest test there is: paste the result into any checker and it comes back flagged anyway, because the text gets normalised before it is scored. One minute with a checker settles more of these debates than any forum thread.

  • Synonym swapping. Turning 'utilize' into 'use' changes one word and leaves the sentence at the same length, in the same order, with the same rhythm.
  • Sending the draft back through a chatbot with an instruction to make it sound human. The rewrite is produced the same way the original was, so it lands in similar territory unless the structure changes.
  • Inserting invisible characters or lookalike Unicode letters. It breaks copy and paste, breaks screen readers, and gets stripped by any pipeline that normalises text.
  • Rewriting only the opening paragraph. The statistics are computed across everything you submit, and the untouched pages carry most of the weight.
  • Padding with filler to dilute the flagged sections. Filler is written the same way the rest of the draft was, so it dilutes nothing.
  • Changing formatting, fonts or spacing. None of it survives the extraction step that turns your file into plain text.
  • Mixing several AI tools, one paragraph each. Different models produce output that shares the same statistical shape, so a blended draft keeps the shape of machine writing.

Deliberate errors are the worst trade on the list

The most repeated tip is to add typos, splice a few commas, drop an article here and there. It can shift the reading, for a reason that should stop you. It raises average surprise by making the writing worse. Every error you plant is a mark you give up with the human being who reads the work and decides what it is worth. You are trading a number you cannot see for a grade you can. Some guides now dress the same idea up as 'controlled imperfections'. It is the same trade in nicer packaging.

There is a second cost. Slightly-off English is the signature of a second-language writer, and that is the group the Stanford figures show being falsely flagged. Broken prose does not read as human. It reads as careless.

A workable rule: do not make an edit you would not defend out loud to the person grading the work. Everything below passes that test, because a good editor would ask for it anyway.

The levers that actually move a score

Sentence-length variance is the single biggest lever. Model output clusters. Left alone it produces sentence after sentence of similar length, each one built as subject, verb, object, qualifier. Human writing does not hold a band. A paragraph might run four words, then thirty-one, then nine. Count the words in each sentence of a paragraph you are worried about. If every count lands within a few of every other, fix that before you touch a single word choice. Three edits do most of the job: split a long sentence at its conjunction and let one half stand alone as a short one, fold two related short sentences into a single longer one with a subordinate clause, and delete the sentence that restates the sentence before it, which a model adds by habit and which you will find in almost every paragraph it writes. Say three sentences run 21, 24 and 20 words. After those edits they run 6, 33 and 11. No error was introduced and no claim changed. Read it aloud and you will hear the difference before any tool tells you.

The second lever is deletion. Models fill slots, and some of those slots produce phrases that carry no information. Removing them cuts the runs of highly predictable words and tightens the prose at the same time. The version without them is better writing. The registry on this site names 44 of these tells and highlights the exact phrases where they occur in your own text, free and unlimited. The categories worth knowing by heart are short.

Above the sentence there is a third pattern most people never touch. Model paragraphs come out at roughly equal length, each with a topic sentence, two supporting sentences and a closer, and the sections come out equal too. Real writing is lumpy because evidence is lumpy. The point with three sources gets a long paragraph, and the point that needs one line gets one line. Put the headings where the argument turns rather than at even intervals.

  • Scene-setting openers that describe an era or an industry instead of the subject. Start at the claim.
  • Dead transitions. A formal additive connector at the head of a sentence that promises more evidence. If the sentence does not add evidence, cut both.
  • The negation pivot, as in 'It is not just X, it is Y.' A model produces this constantly. Once you notice it you cannot stop seeing it.
  • Three-item lists where the third item exists only to make three. Keep the two that carry weight.
  • Stacked hedges. Two qualifiers in one clause say less than one does.
  • The closing sentence that summarises the paragraph you have just read. Almost always deletable.

Measure, edit, measure again

Then measure, because guessing is what makes this slow. Paste 65 words or more into the checker on this page and you get a score from 0 to 100 in seconds, along with one of three readings: reads as AI-generated, borderline, or reads as human-written. A check costs one credit at any length, so there is no reason to split a passage that already fits.

Change one thing, run it again, and you learn which of your own habits is carrying the weight. That feedback loop is the real skill, and it transfers to everything you write next. The page on what a score means explains the bands, and the tells registry shows you the exact phrases to cut before you spend a single credit.

If you would rather start from a rewrite, the humanizer restructures every sentence in one click, up to 12,000 characters per run, which is roughly 2,000 words, and longer documents go through in consecutive runs. It costs one credit for every 50 words. The free tier is 10 credits a day with no account, no card and no email address, and the tell registry, the proofreader, the word counter and the readability check stay free and unlimited. For the version of this work that starts from an honest no-guarantee position rather than a promise, the undetectable page is the companion read.

Questions people actually ask

How to bypass AI detection: what actually changes the reading: common questions

Can you actually bypass AI detection?

Nobody can promise you a particular result from a particular detector, and a tool that promises one is selling something it cannot deliver. What you can do is change the properties detectors measure, which are predictability and its variance. Sentence-length variance, deleted stock phrasing and uneven paragraph structure move those properties. Word swaps and typos mostly do not.

Do AI word swappers and paraphrasers work?

A paraphraser that keeps the sentence structure and trades words for synonyms changes very little, because the reading is computed over sequences and rhythm rather than over individual word choices. It also tends to make the prose slightly worse, since the second-choice synonym is usually the less precise one. Changing the structure does more than changing the surface.

Should I add typos or grammar mistakes on purpose?

No. The effect is small and the cost is not. You are handing away marks with the human reader in order to move a number you cannot verify, and you are making your writing look like the second-language prose that the Stanford study found being falsely flagged. Fix the structure and leave the grammar alone.

Why do two detectors give me different scores?

Different detectors are built on different training data and weigh different features, so their numbers do not line up and no single score is a verdict. What holds across them is the underlying property. Text that is uniformly predictable reads as machine-made, and text with real variation in sentence length and rhythm does not. That is the property worth working on, since it is the one every classifier of this kind is looking at in some form.

What if I wrote it myself and a detector flagged it?

That is a different problem with its own answer. Detectors flag clean, even, correct prose, which is exactly what careful writers and second-language writers produce. The Stanford HAI study found 61% of TOEFL essays by non-native speakers falsely flagged across seven detectors. Start with the false positives page and the proof of human writing tools, which cover what evidence you can assemble and how to talk about it.

Keep reading