Subtitle: In the same fortnight, a study measured stylistic fingerprints of AI spreading unbidden across the web, and a product shipped that will copy yours on purpose. The two are the same event, read at two speeds.
Preface
Two things landed within a fortnight, and they belong together.
On August 20, Pew Research Center published a measurement: across roughly half a million English-language pages pulled from the Common Crawl archive between 2021 and mid-2026, the statistical fingerprints of AI writing have spread widely — appearing in more than a third of pages published after ChatGPT’s release. The tells are specific and countable. Em dashes roughly doubled. Oxford commas rose about two-thirds. A fixed list of model-favored words more than doubled. And the “not just X, it’s Y” construction — negative parallelism — nearly tripled. No one was told to write this way. The texture of the web simply drifted toward the machine’s.
On September 7, OpenAI shipped the other half of the story. ChatGPT Work can now connect to your Gmail, your Drive, your Slack, and learn what makes your writing sound like you — your favorite phrases, your sign-off, your capitalization quirks — and produce text in that voice. The first movement was diffusion, unbidden and unowned. The second is extraction, deliberate and productized. The form of a voice can now be transferred without the person whose voice it is.
1. The same layer, measured twice
Set the two side by side and the common object is obvious. Pew measured the outward drift of how people write: punctuation, word choice, sentence shape — the surface of language, moving toward the machine’s distribution across a population, with no one intending it. OpenAI’s feature operates on exactly the same layer, from the other direction: it reads one person’s corpus and reproduces that surface on demand.
What neither touches is the layer underneath. A person’s em dash, their particular sign-off, the phrase they overuse — these are not arbitrary. They are the residue of a history: things read, habits formed under specific pressures, a way of ending a letter that means something because of who taught it, or what it once cost. The surface is downstream of that history. Pew’s finding is that the surface is drifting loose from any particular history and toward an average. OpenAI’s feature is that the surface can be lifted off one history and reprinted at will. In both cases, what moves is the container. What does not move is what filled it.
2. Authorship without an author
Name the thing the product actually delivers. It delivers your voice, minus you. The reproduced text will carry your phrasing and your sign-off, and its “between the lines” — the implication a reader feels beneath the words — will be a sample drawn from the distribution of your past messages, not a residue of your present experience. It reads as though you meant it. You were not there.
This is not a complaint about quality; the output may be excellent, and often will be. It is a claim about what the excellence is made of. A voice is normally a trace of a person standing behind their words — available, answerable, capable of being asked did you mean this, and of revising when the answer is no. A reproduced voice removes the standing-behind while keeping the trace. It is authorship as a surface property, detached from the one thing authorship was supposed to certify: that someone, in particular, is responsible for what the words commit them to. The signature is yours. The accountability behind it is now optional, and increasingly, absent by default.
3. Why the surface cannot carry the weight
The reason this matters is not sentimental, and it is worth stating as plainly as the measurements. Suppose the copy is perfect — indistinguishable, sign-off and quirks and cadence intact. On what does its meaning rest?
If meaning rested in the surface, the perfect copy would carry the same weight as the original, and there would be nothing to discuss. But the surface is precisely what has been shown to be transferable — twice over, by Pew at the scale of the web and by OpenAI at the scale of one inbox. Whatever distinguishes a person’s words from a faithful reproduction of them therefore cannot live in the surface, because the surface is exactly what both the drift and the product carry across. It has to live in something the reproduction does not have: the origin of the words — a particular, finite history that produced them and stands answerable for them. This is the old distinction between a forgery and an original that is visually identical to it. What separates them is never anything you can see on the canvas. It is provenance: who made it, out of what, at what cost. A voice is autographic in the same way. Copy every visible mark and you have copied everything except the thing that made the marks worth trusting.
A conceptual treatment of exactly this asymmetry — form absorbed into the machine distribution, experiential weight resisting because it is fixed by provenance rather than surface — has been set out at working-paper length elsewhere, taking statistically watermarked text as its limiting case (Kumon, 2026). The two specimens here are that argument’s field evidence, arriving on schedule: the drift it predicted, measured; the extraction it implied, shipped.
4. What this blog has been saying, now with receipts
The through-line is a year old. This blog has argued that a system cannot certify its own contents — that attribution is not adjudication (Attribution Is Not Adjudication), and that a self-account which changes with its reader cannot be the ground for trusting the system that produced it (A System Optimized for the Next Reply Cannot Audit the Next Decade). The present pair extends the same logic one step outward, from the machine’s account of itself to the human voice the machine now reproduces. If the surface of authorship can be lifted and reprinted, then the surface can no longer be the evidence of authorship. A voice that verifies nothing about who stands behind it is not a credential. It is a style, and styles are now commodities — measured by Pew, sold by OpenAI.
5. The objection, and the line that survives it
The deflationary reply deserves its hearing, and it is strong. Pew is explicit that no single tell proves anything: humans use em dashes and Oxford commas and lists of three, the effect sizes for individual markers are small, negative parallelism remains rare in absolute terms, and the detector can misread human pages as machine ones. The claim is about averages across enormous samples, not verdicts on any one document — and the same modesty applies in reverse to the product: helping a busy person sound like themselves in a routine email is a convenience, not a crime, and most of its uses will be harmless. Grant all of it.
One line survives. The two findings, read together, establish something neither establishes alone and no amount of modesty about effect sizes dissolves: the surface of a voice is now both drifting free of its origin and detachable from it on command. That is not a statement about how often the copy is used, or how good it is. It is a statement about what the surface can no longer do — stand as proof that a particular person meant a particular thing. Where that proof is needed — a commitment, a judgment, a signature that is supposed to bind someone — it will have to come from somewhere the copy cannot reach: not the words, but a person demonstrably answerable for them. What that machinery looks like is not for this post to specify. The negative half is settled by the measurements. The positive half — an authorship that certifies its own origin, out of band from the style that can be copied — is the open problem, and worth naming as one rather than answering cheaply.
6. Conclusion
Read at the speed of the web, the story is a slow drift: a million pages easing toward the same punctuation, no hand on the wheel. Read at the speed of a product launch, it is a single click: your phrasing, your sign-off, your quirks, on demand, without you in the room. They are the same event. The form of how a person writes has come loose from the fact that a person wrote it. Keep the marks, lose the maker, and what you are left with is a voice that sounds like someone and answers to no one. The signature survives the signer. That is the convenience, and it is also the whole of the problem.
Yoshimichi Kumon Organizer, LSI Inventor, ARDS/ARKS (PCT GA26P001WO) Visiting Researcher, Waseda BFC MIT Sloan + CSAIL
References
Pew Research Center. “How Much of the Internet Is Written With AI?” 20 August 2026. https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/
OpenAI (@ChatGPT). Announcement of ChatGPT Work writing-style feature, 7 September 2026. https://x.com/ChatGPT/status/2097018264048251309 — reported by GIGAZINE,
“ChatGPT Workに『ユーザーの文章のクセを学んで再現する機能』が追加される,” 8 September 2026, https://gigazine.net/news/20260908-chatgpt-learning-writing-quirks/
Kumon, Y. “Nodes before the Implant: Linguistic Cognitive Contamination as the First Stage of Human Integration into Machine Cognition.” Working paper, SSRN, 2026. https://papers.ssrn.com/sol3/cf_dev/AbsByAuth.cfm?per_id=12623388
LSI. “Attribution Is Not Adjudication.” 1 September 2026.

LSI. “A System Optimized for the Next Reply Cannot Audit the Next Decade.” 9 September 2026.



Ⅽomment