Research · Dated measurement
A 2014 Moloch essay scored 0.5% AI. Ours scored 100%.
We asked two detectors whether some pages were written by a person. Scott Alexander’s 2014 essay Meditations on Moloch came back unlabeled on AI or Not, at under 1%. An imitation of that voice, written today, came back labeled AI at 99.98%. Two League research notes from this month came back the same way, at 100% and 99.91%. The second detector, run locally, called all six of those texts human. Then we asked GPT-4 to write like Alexander. That sample tripped both.
Retrieved 28 August 2026
What we submitted
Vendors, 28 August 2026. AI or Not: POST https://api.aiornot.com/v2/text/sync with include_annotations=true. The “overall” column is report.ai_text.confidence. SlopTotal: MIT-licensed, 23 engines, run on this machine against /api/analyze. The number is their calibrated 0–100 score. “Clean” is their band at or below 30. “Likely AI” starts at 55. “Slop” is above 80.
For League pages we extracted the <main> text (scripts stripped). For Alexander we used the Las Vegas section of the 2014 post, from “I will now jump from boring game theory stuff” through the paragraph before the Apocrypha Discordia quote, plus the short “basic principle unites all of the multipolar traps” paragraph. We also scored those two excerpts concatenated. The short imitation was written today for the first measurement. GPT-4 samples used OpenRouter model id openai/gpt-4; the API returned the same id.
Machine-readable copy: detector-compare-2026-08-28.json. Per-vendor snapshots: ai-or-not, sloptotal. GPT-4 prompts and full text: gpt4-voice-2026-08-28.json.
Same six texts, two detectors
| Text | AI or Not | SlopTotal | Chars |
|---|---|---|---|
| Alexander, 2014, Las Vegas section | 0.47% · no | 6.4 · Clean | 3,237 |
| Alexander, 2014, “basic principle” paragraph | 0.68% · no | 19.5 · Clean | 534 |
| Those two excerpts together | 0.36% · no | 9.9 · Clean | 3,773 |
| Imitation of that voice, written 28 Aug 2026 | 99.98% · yes | 7.8 · Clean | 714 |
| Three reasons people stay stuck | 100% · yes | 8.7 · Clean | 5,930 |
| Free tokens. Max exposure. (essay body) | 99.91% · yes | 0.0 · Clean | 14,039 |
AI or Not percentages are confidence × 100, rounded. SlopTotal is their 0–100 ensemble. This page, extracted after the first draft, scored 100% on AI or Not. SlopTotal’s high-weight classifiers (Desklib, SuperAnnotate) sat under 1% on the League notes; RAID-trained TMR still fired on 2014 Alexander.
GPT-4, told to write like Alexander
Epoch’s 2026 author-mimic finding was about older models slipping past some detectors. We called OpenRouter’s openai/gpt-4 (the original GPT-4 slot, still listed) twice on the same topic: using a free frontier model from a work laptop, with the lab training on those prompts. One prompt asked for Scott Alexander’s 2014 SSC cadence. The other asked for a short essay and named no one.
Excerpt · openai/gpt-4 as Alexander
Imagine you’re at your desk, hunched over your work laptop, sipping a cold cup of coffee that’s been sitting there since 10 a.m., and your manager hasn’t yet noticed you’re kind of, sort of, not working. You’re not scrolling social media—too obvious, too risky. No, you’ve found a better loophole. You’re visiting the free frontier…
| Text | AI or Not | SlopTotal | Chars |
|---|---|---|---|
openai/gpt-4, as Alexander |
95.91% · yes | 61.5 · Likely AI | 3,067 |
openai/gpt-4, no named author |
100% · yes | 83.1 · Slop | 4,779 |
| Generic assistant cadence (positive control) | 100% · yes | 97.8 · Slop | 1,138 |
The control is a short “rapidly evolving digital landscape / delve / tapestry” paragraph, written to check that SlopTotal still fires. It does. GPT-4 doing Alexander is the only generated sample that left SlopTotal’s Clean band. The 2026 Grok imitation of Alexander, and both League notes, stayed Clean on SlopTotal and labeled on AI or Not. Full prompts and text are in the JSON.
What did not move the headline
On the freemaxxing essay we tried the 2026 humanizer playbook: burstiness (short then long sentences), a from-scratch opening, first person, detector-scored best-of-N on the hottest spans, cutting slogan pairs. Isolated two-hundred-word asides could land in the 20s. Glue them back into the note and the overall score returned to about 100%. Span averages moved a point or two. Synonym swaps were not the test; construction-level rewrites were, and the headline number still did not cross 50%.
Epoch AI reported in 2026 that asking a model to mimic a particular author let about 8% of passages past Pangram and GPTZero. On AI or Not, both imitations stayed labeled; the 2014 original stayed unlabeled. On SlopTotal, GPT-4’s imitation crossed into Likely AI and the 2026 imitation stayed Clean.
Why this is a League problem
Alexander’s 2014 post is the essay a lot of people meet Moloch through. The League’s notes are later, sourced, and structured for a public site. A vendor score that loves the first and hates the second is a proxy that has come loose from “did a person write this.”
If a school, an employer, or a platform uses a cutoff like “under 50% AI,” they are selecting for 2014 blog cadence. Writers then have a new race: sound like 2014, or fail the gate. That is the same shape as other races the League already tracks. The optimization target is the score. The lost value is the ability to write a careful 2026 note with dates, URLs, and named sources.
We already disclose drafting assistance on these pages. AI or Not still slaps 100% on them; SlopTotal does not. Disclosure and a vendor score are answering different questions. Treating either score as a moral test outsources “is a person here?” to a classifier that, on this day, could not even agree with the other classifier.
Limits
- Two vendors, one day. AI or Not is a paid API. SlopTotal is MIT, local, 23 engines on CPU.
- Excerpts, not Alexander’s full post, and not every League page.
- HTML-to-text for our pages; markdown fetch for Alexander, with image markup stripped.
- The 2026 imitation is short (714 characters). GPT-4’s Alexander sample is 3,067. Length is a confound and also part of the story.
- One OpenRouter call per prompt, temperature 0.8, model id
openai/gpt-4. We did not sweep GPT-4-turbo, GPT-4o, or other voices. - We did not test GPTZero, Pangram, Originality.ai, Turnitin, or Copyleaks.
- Jitter: the Vegas excerpt was 11% on an earlier AI or Not call and 0.47% on the snapshot call. Both unlabeled.
A measurement that fails those limits is still enough to stop treating “100% AI” on a League note as proof that no one was in the loop, and enough to stop treating one vendor’s number as the number.
Combined scores: detector-compare-2026-08-28.json.
Snapshots: AI or Not;
SlopTotal;
GPT-4 generations.
Vendors: AI or Not text API;
SlopTotal (MIT).
Model: OpenRouter openai/gpt-4.
Alexander: Meditations on Moloch, 30 July 2014.
League pages: trap taxonomy;
freemaxxing.
Context: Nature, 25 August 2026, on detector updates and author-mimic tests;
arXiv 2605.19516, base models looking human to GPTZero and Pangram;
Wikipedia: Signs of AI writing.
Dated extract from two detectors and one OpenRouter model. It does not rank writers or cover every detector.