The Slop Score

How much of the writing is made of familiar AI patterns?

A public 0–100 benchmark of words and constructions that language models overuse. Lower is better.

Public leaderboard snapshot

Slop Score

November 7, 2025 · lower is better

  • Human baseline10.4
    83 samples
  • Kimi K2 090518.3
    300 samples
  • Claude Sonnet 4.519.5
    150 samples
  • GPT-5 mini26.3
    300 samples
  • ChatGPT 4o47.7
    300 samples
  • Gemini 2.5 Flash77.6
    300 samples
Public benchmark snapshot dated November 7, 2025. Scores are calculated from standardized creative-writing and essay prompts.

Measurement

Three signals make the score

60%Overused words
Listed-word matches per 1,000 words
25%Contrast scaffolds
Repeated “not X, but Y” constructions per 1,000 characters
15%Overused three-word phrases
Listed phrase matches per 1,000 words

Scores are normalized against the benchmark cohort and can shift as that cohort changes. The benchmark uses 300 standardized prompts: 150 creative-writing tasks and 150 essays. Sample counts vary in this public snapshot.

Read the public methodology (opens in a new tab)