Across kinds of writing
Each cell is one kind of writing, with the verdict of the model that kind is decided against, named in its column head, against the people of that kind.
▲ model marker the model uses it more · ▼ more human the person uses it more · ◆ careful writing both, more than casual writers · – no signal · · not recorded no text of that kind can show it
Not covered: email, chat, product reviews, and social platforms other than Reddit, such as X, Facebook and LinkedIn. Student essays are covered, with the machine side written for this project rather than taken from a published set.
The markers in research abstracts
Every column up to the last model is the same set of documents: written by a person, and written again by each model from the same prompt (RAID). Click a row to see what was counted.
One document, every writer
The same prompt answered by every writer. Counted markers are highlighted. Documents are drawn at random from the ones every writer covered, not chosen.
Check a text
Paste anything. The markers in it are highlighted and each one is shown next to how often people use it in the kind of writing picked above. There is no score, because the tables above do not support one. Nothing leaves your browser.
How it is counted
- One document or one assignment, several writers, per kind of writing. Three kinds of writing are measured. Research abstracts and Reddit posts are each a set of documents written before ChatGPT existed, by a person as far as the source can tell (the Reddit file was not filtered for bots or spam), and the same documents written by GPT-3.5, GPT-4, Llama 2 chat and Mistral chat from the same prompt, the document’s title (RAID, MIT; the 2023 model checkpoints, greedy decoding with no repetition penalty). School essays share no document: a class answered one of seven assignments, and Claude and Llama 3 were given the same assignment and wrote essays of their own, for this project, in September 2026. Each kind is measured on its own, with its own person, placebo and correction.
- Two measures, two verdicts. A word or phrase is counted per thousand words over every text. A phrase used once in a short text has a higher rate than the same phrase in a long one, and the models write to different lengths than the people, so the verdict compares a model with the person only between texts of about the same length, and only on the documents the model still has after the checks below (the rates shown are the whole columns). The test allows for a writer who repeats a word within one text, and it has to survive a correction for testing every marker at once. A property of the whole text, such as no contractions anywhere, is the share of texts that have it, and its verdict comes from those shares.
- One verdict, and the count across that kind’s models. Each kind of writing is decided against one model, the one named in its column head, and the table’s verdict compares that model with the person, as it always has. The grid adds how many of that kind’s own models the marker separates from the person there, each model decided by exactly the same rules. How many there are is a property of the kind, not a fixed four: four where a published corpus had four models write the same documents, two where two models answered the same assignment, and the column head says which. A model that does not separate the marker is counted only where there was something to compare: at least five uses of a word between it and the person, or, for a property of the whole text, thirty pairs and five texts on the property’s rarer side. The count describes; it is not a further test.
- Pairs, not file order. For a share, each model’s text is paired with the person’s text for the same document, and a pair is kept only when both fall in the same length band. In school essays nobody wrote the same document, so a model’s essay is paired with a student’s essay written to the same assignment and in the same length band, drawn at random within both; those pairs are looser than a pair on one document. The comparison columns share no documents with the person, so they are matched by length from a random draw within each band. A text too short for a marker to say anything about, such as one with fewer than five sentences for every sentence the same length, is left out of that marker’s share instead of counted as a no.
- Patterns, shown. Every marker is a regular expression. For every cell the page can show what it matched: the forms, the word in front, and five sentences picked by a seeded shuffle. Sentences are quoted from RAID’s model texts (MIT), from arXiv abstracts (CC0) and from the columns written for this project. The people’s Reddit posts are counted and never quoted, and neither is a model’s sentence that repeats five words in a row of the person’s post, or eight of anyone’s. The students’ essays are counted and never quoted, and they are not in this repository; a model’s sentence that repeats eight words in a row of any student’s essay is not quoted, and a model’s essay is shown whole only when no eight words in a row of it were written by exactly one student. Casual and careful writing are linked to where they were posted, not quoted.
- A placebo. The person’s texts in each kind of writing are split at random and the same tests run on both halves. Every row should tie; a row that does not is marked. The class in school essays is far larger than any model’s column, so there the two halves are cut to the size of the model’s own comparison, assignment by assignment and length band by length band.
- Dropped before counting. A model text that stops mid-sentence at its length limit, that is a refusal or a label instead of the text asked for, or that shares more than half its five-word sequences with the person’s document is dropped from its column; the person’s texts are never dropped for these. A document is left out of every column, the person’s included, when any of its writers wrote it in a language other than English. A research abstract with an arXiv version posted between ChatGPT’s release and the date of RAID’s file is left out of every column once its paper has been dated, since its text may have been revised after ChatGPT. In school essays a model’s essay counts as remembered when more than half its five-word sequences occur in the students’ essays to the same assignment, since there is no one document to check it against; every student essay measured was released in the Kaggle Feedback Prize competition of December 2021; and the corpus had already replaced the students’ names with placeholders, which are taken out before counting, while the models’ letters are left as written, the names they made up included. The columns generated through Claude Code, in abstracts and in essays, are reported by the cut-off and not-an-answer checks but not cut by them.
- Casual and careful writing. Casual writing is everyday online comments; careful writing is edited question-and-answer posts; both are from before ChatGPT existed, so a person wrote them (sources: Hacker News and Stack Exchange). The third comparison column is GPT-3.5 answering questions (HC3). These are other kinds of writing, so they are not tinted; they show how much of a marker is about the kind of text rather than the writer.
- Not covered. Email, chat, product reviews, and social platforms other than Reddit (X, Facebook, LinkedIn) are not measured here: no set that pairs a person’s text in those with models writing the same thing can be used here under its terms. Student essays are now covered, with the machine side written for this project rather than taken from a published set; the note under that kind’s table says by what and when.