Skip to content

Glossary

AI content detection

Tools that estimate whether writing was produced by a machine, returning a likelihood rather than the evidence people treat it as.

In plain terms

Software that guesses whether a person or a machine wrote something. The guess arrives looking like a verdict, which is the whole problem: people act on it as though it were proof, and the person on the receiving end has no way to show it is wrong.

01

Why it matters

Because these tools get pointed at people rather than at content. A student, a candidate or an employee gets accused on the strength of a score, and the accusation is unusually hard to answer, since nobody can produce evidence that they thought of something themselves.

02

How it works

The output is a likelihood, not a finding. A detector reports how closely a piece of writing resembles patterns it associates with machine generation, which is a statistical statement about text, and it is read as a factual statement about a person.

The errors fall unevenly. A missed detection costs almost nothing, and a false positive lands on somebody as an accusation of dishonesty, so the two kinds of mistake are not equivalent even at identical rates.

There is no way to disprove one. Asked to demonstrate that they wrote something themselves, an honest person has nothing to offer beyond saying so, which means the tool creates an allegation the accused structurally cannot answer.

Accuracy claims come from whoever is selling the detector, measured on material they chose. That is not necessarily dishonest and it is not independent, and the conditions where a detector was measured are rarely the conditions where it gets used.

Detection chases generation and cannot catch up. Every detector is trained on the output of models that already exist, and writing produced by newer ones, or run through an ordinary editing pass, moves away from whatever the detector learned.

Detecting an image or a recording is a different problem from detecting text, and the two get discussed as one. Generated media can carry signals embedded at the point of creation, which is a question about provenance; a passage of prose carries nothing but itself, which is why text detection is guesswork in a way the others need not be.

Careful human writing resembles machine writing. Clear structure, even sentence lengths and unremarkable vocabulary are what a competent writer produces under instruction, and they are also what these tools score as suspicious, which puts the most disciplined writers at the greatest risk.

What the tool says, and what people hear

What the tool says, and what people hearThe translation between the columns happens for an ordinary reason rather than a careless one: a decision is binary and a score is not, so somebody has to convert one into the other, and no threshold makes that conversion honest. Set it high and the tool catches almost nothing worth catching. Set it low and it starts accusing people, mostly the ones whose writing is most disciplined. There is no setting where it behaves like the evidence it is being asked to be. What makes this worth a firm position rather than a caution is the asymmetry of the harm: the person wrongly flagged carries the whole cost, cannot produce anything that would clear them, and is usually facing an institution that believes it has a technical basis for what it is saying. If the question genuinely matters, ask for the process. Drafts, revision history and a five-minute conversation about the work answer it directly, they are checkable by both sides, and they leave the person something to say.What it reportsThis text resembles generatedwriting.To this degree of confidence.Against patterns from knownmodels.What gets acted onThis person used AI.We are sure enough to act.The result is a finding.The left-hand column is whatthese tools are entitled toclaim, and every line of it ishedged. The right-hand column iswhat happens once a numberreaches somebody who has to makea decision, and none of thehedging survives the trip.
The translation between the columns happens for an ordinary reason rather than a careless one: a decision is binary and a score is not, so somebody has to convert one into the other, and no threshold makes that conversion honest. Set it high and the tool catches almost nothing worth catching. Set it low and it starts accusing people, mostly the ones whose writing is most disciplined. There is no setting where it behaves like the evidence it is being asked to be. What makes this worth a firm position rather than a caution is the asymmetry of the harm: the person wrongly flagged carries the whole cost, cannot produce anything that would clear them, and is usually facing an institution that believes it has a technical basis for what it is saying. If the question genuinely matters, ask for the process. Drafts, revision history and a five-minute conversation about the work answer it directly, they are checkable by both sides, and they leave the person something to say.
03

Seen in the wild

  • A submission scored as machine-written with nothing the author can offer in reply.

    ChatGPT
  • An editing pass changing the score without changing who wrote the piece.

    Grammarly
  • Drafts and revision history answering the question the detector cannot.

    Notion AI
04

Common misconceptions

People assume

A high score is evidence somebody used AI.

In fact

It is an estimate of resemblance to patterns the tool associates with generated text. Treating that as evidence converts a statistical statement about writing into an accusation about a person, which it cannot support.

People assume

Better detectors will solve this.

In fact

Detection is always trained on models that already exist, so it trails generation by design. Improvement narrows the gap in a moment and does not close it, and an ordinary editing pass reopens it either way.

05

Questions

Can we use one to check submissions?
Not as a basis for a decision about a person. The output is a likelihood rather than evidence, the accused has no way to disprove it, and the cost of being wrong falls entirely on them. As a prompt for a conversation it is defensible; as a finding it is not.
Why do careful writers get flagged?
Because the qualities these tools score as machine-like are the qualities of disciplined writing: clear structure, consistent sentence length and plain vocabulary. Somebody following a style guide closely is producing exactly the signal the detector was built to notice, which puts the most careful writers in the most exposed position.
What should we use instead?
Evidence about the process rather than the artefact. Drafts, revision history and a short conversation about the work answer the question a detector only estimates, and they produce something the person can actually engage with rather than an unanswerable score.
06

Key takeaways

  • The output is a likelihood about text, read as a verdict about a person.
  • False positives cost far more than misses, so equal rates are not equal.
  • Nobody can disprove one, which makes it an unanswerable accusation.
  • Process evidence beats detection: drafts and history are checkable.
08

Tools that use this

  • ChatGPT

    A submission scored with nothing the author can offer in reply.

  • Grammarly

    An editing pass changing the score, not the authorship.

  • Notion AI

    Drafts and revision history answering what a score cannot.

Last checked August 2026

All glossary terms