How to Remove AI Slop From Your College Writing
A practical guide to replacing generic AI prose with specific, defensible writing that preserves your own ideas and voice.
At a glance
- Slop is generic, ungrounded writing—not a synonym for every use of AI.
- Surface-level humanizers can make a weak draft less trustworthy instead of restoring the writer's voice.
- The durable remedy is to develop and support the argument with the writer's own thinking and evidence.
The lone way to effectively remove AI slop is to generate the text you turn in yourself, out of your own thinking and ideas, rather than trying to mask an AI draft.
What "AI Slop" Actually Means
Three dictionaries landed on the same word at the same time. The Macquarie Dictionary made "AI slop" its 2025 Word of the Year, defining it as low-quality content created by generative AI, often containing errors, and not requested by the user. Merriam-Webster chose "slop" for 2025 as well: digital content of low quality that is produced usually in quantity by means of artificial intelligence. The American Dialect Society voted the same way on January 9, 2026.
The term was popularized by a May 2024 blog post from web developer Simon Willison, who argued that just as "spam" became the word for unwanted email, "slop" is the word for unwanted AI-generated content. Willison was careful about the boundary: "Not all promotional content is spam, and not all AI-generated content is slop." What makes it slop is that nobody reviewed it and nobody asked for it. The classic student version is a paper copied straight out of ChatGPT, unedited and unchecked.
Slop is not a synonym for AI; it is a name for a lazy, thoughtless use of those tools that easily gets detected by machines and loathed by readers.
Why Your ChatGPT Draft Reads Like Slop
The first reason is a very small set of tell-tale words. A computational-linguistics study by Tom Juzek and Zina Ward tracked 21 of them across PubMed abstracts and compared their rate per million words in 2020 against 2024: "delve" rose about 1,375 percent, "underscores" about 904 percent, "intricate" about 611 percent, and "realm" about 381 percent. A separate analysis of more than 15 million PubMed abstracts estimated that at least 13.5% of 2024 abstracts had been processed with a language model, reaching 40% in some subfields. Words like these are a stamp of artificiality.
Beyond vocabulary, slop has a shape that is instantly recognizable. The language begins formulaically, with phrases like "it is important to note." It closes, similarly, with "in conclusion." In between we get paragraphs of pale generality, the kind of thing that obviously has to be true but usually goes unsaid. The sentences have a flat rhythm, almost uniform, with scant variation.
But there is a deeper reason. It is about what the writing commits to. The language sounds confident, and yet for every specific claim, example, or piece of evidence, a model will happily supply a paragraph's worth of filler that commits to nothing. Professors read it as slop precisely because it refuses to commit.
Why AI Detectors Can Spot Slop in the First Place
The Legacy Signals: Perplexity and Burstiness
In the earliest years of AI detection, perplexity and burstiness were the main signals statistics-based detectors used. Perplexity measures how predictable each word is given the surrounding context: if a sentence can be mostly guessed from what surrounds it, it has low perplexity. Language models are built to do exactly that, so they produce low-perplexity text. Burstiness measures variation — how much sentence length and grammatical structure fluctuate across a document. Human writing tends to be bursty; model output tends to be flat.
Why the Legacy Signals Aren't Enough
Those two numbers are too blunt to use on real student work. Pangram's researchers argue they cannot reliably detect AI writing while keeping a false positive rate low enough for production: famous human-written texts that sit in model training data get flagged, different models have different perplexity signatures, closed commercial models don't expose the probabilities the method needs, and non-native English speakers draw elevated false positives. Sprinkling random variation into a model-written sentence might fool a tired human reader. It is not what a modern detector is looking at.
The New Guard: Neural Classifiers
Pangram is an example of the newer generation: a classifier built on the same kind of language-model architecture that writes AI text in the first place, trained on roughly 1 million documents to separate human-authored from machine-generated writing. What it picks up are structural fingerprints — coherence patterns and logical flow across a document's paragraphs and sentences.
Why Surface Edits Fail
Because the signal lives at the structural level rather than at the level of word choice, swapping synonyms or tweaking the surface of a sentence barely dents it. Serious detectors analyze the overall architecture of an argument, not isolated sentences. Removing AI slop therefore has to go deeper than tweaking and slashing at the surface.
The Real Stakes: Turnitin and the Fear of a False Flag
Two Turnitin numbers get quoted, and they measure different things. The company claims its tool is 98% accurate at identifying AI-generated content. That figure is about the text it does flag. It says nothing about how much AI writing slips past.
Turnitin's chief product officer, Annie Chechitelli, gave BestColleges the other half: "We would rather miss some AI writing than have a higher false positive rate. So we are estimating that we find about 85% of it. We let probably 15% go by in order to reduce our false positives to less than 1 percent." So roughly one in seven AI-written passages goes uncaught, on purpose, to keep innocent students from being flagged. By June 2023, Turnitin had also admitted that real-world false positives were running higher than its lab results.
Independent testing was worse. Washington Post reporters ran 16 sample documents through the detector in April 2023 and found it got over half of them at least partly wrong, including flagging part of a fully human-written high school essay.
OpenAI's own attempt fared no better. It released a classifier in January 2023, trained on AI-generated text against human-written text, and pulled it on July 20, 2023, citing poor accuracy: it caught only 26% of AI-generated text while wrongly flagging 9% of human writing. Even the people who build the models could not reliably tell the two apart.
The False-positive Trap That Punishes Innocent Students
When Simple Writing Looks Like a Machine
A 2023 study in Patterns by researchers at Stanford put seven detectors up against 91 TOEFL essays written by non-native English speakers and 88 essays by US eighth-graders. Every essay was human-written. The average false-positive rate on the non-native essays was 61.3%; on the native-speaker essays it was 5.1%. Same task, same honesty, twelve times the risk of being accused. The mechanism is the perplexity problem again: detectors read simpler vocabulary and lower sentence complexity as evidence of a machine, whether the writer chose that vocabulary by design or by limitation.
Universities noticed. In August 2023, Vanderbilt announced it was turning off Turnitin's AI detector, citing the false-positive imbalance and doing the arithmetic on its own submissions: "Vanderbilt submitted 75,000 papers to Turnitin in 2022. If this AI detection tool was available then, around 750 student papers could have been incorrectly labeled as having some of it written by AI." By September, Michigan State, Northwestern, and the University of Texas at Austin had also declined to use it.
The Real Human Cost
Marley Stevens, a junior at the University of North Georgia, is the case students keep citing. She got a zero on a criminal justice paper after Turnitin flagged it, having used only Grammarly to check her writing. The zero pulled her grade below the 3.0 GPA her HOPE Scholarship required, she lost the scholarship, she was placed on academic probation, and she had to pay $105 for an academic integrity seminar.
What Humanizers Actually Do
The pitch is that you can bypass Turnitin without doing the messy work of rewriting your draft. The most thorough look under the hood is DAMAGE, a January 2025 paper by researchers at Pangram Labs that ran text through 19 humanizer tools. Two families of strategy account for most of them, and a couple of tools add tricks on top.
| Strategy | What the tool does | What goes wrong |
|---|---|---|
| Synonym substitution | Rules-based swaps from a hand-built list of "human-like" alternatives, leaving sentence structure intact | The dictionary has no sense of context, so you get tortured phrases — Pangram's example is "counterfeit consciousness" in place of "artificial intelligence" |
| LLM paraphrase | A language model rewrites the passage for fluency and variety | The paraphrase is only loosely grounded in your text; meaning drifts, and the model can invent whole sentences, including citations to sources that were never there |
| Invisible characters | Homoglyphs — the Cyrillic "а" for the Latin "a" — plus zero-width and thin spaces, to confuse the detector's tokenizer | Detectors now check for non-standard characters and spacing errors, so the tampering itself becomes the signal |
| Deliberate errors | Injected typos and broken grammar, imitating human imperfection | The output is flawed by design; one review of StealthGPT found it producing nonsense like "watered by the Irrawaddy" |
The perverse part is that the better a humanizer works as prose, the worse it works as camouflage. Pangram's August 2025 benchmark puts it plainly: "The more readable/fluent the humanizer text is, the more likely it is to be detected by Pangram." You can have text that reads well, or text that has been mangled enough to look unfamiliar to a classifier. You cannot have both.
None of these tools remove slop. They mask its origin, and they leave you with writing that is sloppier than what you started with and dishonest in its aim.
The Detectors Already Caught Up
There was a window where this worked. A 2023 NeurIPS paper showed that a strong paraphraser, DIPPER, could drop DetectGPT's accuracy from 70.3% to 4.6% at a constant 1% false positive rate, without much changing what the text meant. That result is what created the humanizer market.
The window closed. A University of Chicago working paper by Brian Jabarian and Alex Imas, released in October 2025, found that Pangram's false negative rate stays low even on AI passages run through humanizers such as StealthGPT, while GPTZero's false negative rate ran around 50% and above across most genres and models. Running your draft through a humanizer now buys you the risk without the payoff.
The Real Fix: Rebuild It, Don't Paraphrase It
Picture two students holding the same weak ChatGPT paragraph. One runs it through a paraphraser that swaps a few words and reshuffles some clauses. The other throws the paragraph away and rebuilds the point from their own outline, with their own specific details.
To a good detector these are not close. The paraphrased version still carries the shape, order, and architecture of the original, so it still reads as machine-written. The rebuilt version was assembled the way human thinking assembles an argument, so that is how it reads.
The principle: the fingerprint of AI writing lives in the shape of the discourse, not only in the choice of words. Swapping synonyms in vague prose does nothing. Concrete names, numbers, and examples make writing less sloppy and less detectable at the same time — because they are the same problem.
That principle is also what distribution fine-tuning is built on. The Deft model is a Qwen3-14B model tuned on real human web text from the FineWeb dataset so that it writes fresh prose from an outline instead of paraphrasing model-average text. That is a tool for writers who own and publish what it produces. It is not a way around your school's rules, and it does not do the part of this that matters for a paper with your name on it: the outline and the evidence have to be yours.
How to Actually Do It
Start from Your Own Outline and Evidence
A rough outline is the real seed of writing that is genuinely yours, rather than the generic average a blank prompt produces. Load it with specifics dug out of your own sources before you draft a sentence: names, dates, quotations, examples. A fact-dense style is the best defense against slop, and it is also the reason nobody else's paper will read like yours.
Write the Paragraphs Yourself
Most students start with an AI-drafted paragraph or two and then try to clean it up by paraphrasing. It won't help much. A language model builds a paragraph from an outline and a set of evidence; so must you. You cannot humanize language you did not compose. Reconstructing the paper from your outline and your evidence kills the problem at the root.
Check Every Specific You Kept
It is tempting to keep the details the model produced, because they look plausible. That is the trap: a model can invent a name, a date, a number, or a quotation, and all of them will look plausible. One false citation can sink a paper. Check every specific against the original source. An awkward sentence you wrote yourself is not great prose, but nobody will accuse you of fabricating it.
Read It Aloud, Then Disclose
Read the finished draft aloud once and cut anything that doesn't sound like you or carry something real. Then follow your institution's policy on honestly disclosing what help you had from AI. It is more work than running the draft through a humanizer. It is also the only version of this that produces writing that is genuinely good and actually yours, rather than something that fools a detector.