We went looking for whether it works.
Five pre-registered studies, six hundred-odd generations, four model families. The strongest result first: given a brief and nothing else, 0 of 15 drafts landed in the writer’s target rhythm. Given the same brief and one voice file, 13 of 15 did. Nine turns into a long session it was still landing them, 9 of 16 against 0 of 8 with no file. Everything below is the working, including the two studies that came out against us.
The file steers the model.
The test is pre-registered and it is the same one every time. Take a writer’s profile, take three briefs, and ask a model to write the piece — once with nothing attached, once with the voice file in the system prompt. Then measure whether the sentence rhythm lands in the band that writer’s profile asks for. With no file, 0 of 15 drafts landed in band. With one, 13 of 15 did (p < 0.001), and mean sentence length moved 16.5 words to 9.1 against a target of under twelve.
Then the question underneath it: does it hold nine turns later? Twenty-four ten-turn sessions, with unrelated questions in between to dilute the context the way real work does. At turn 10, drafts written with a voiceprint landed in the writer’s rhythm 9 times out of 16. Drafts with no file landed 0 times out of 8 — and never got close at any point in the session, 2 in band out of 32, sentences averaging 18 words against a target of under twelve. It does not drift, because it never arrives. That replicated on a second model family: 11 of 16 against 0 of 8, p = 0.0001.
How much the file changes depends on how far the model already sits from you, and that is a result rather than an excuse. On Gemini 3.7 Flash, whose untouched sentences average 16.5 words against this writer’s under-twelve target, the file closed nearly all of the gap. On GPT-OSS 120B it moved 6 of 15 to 13. On Claude Sonnet 4.6 there was nothing to move: its default writing already landed in her band 14 times out of 15, so that arm measures the pairing, not the file. It is checkable in about a minute — ask your own model for a paragraph and count.
The file is also, measured across 313 saved drafts on seven of the thirteen dimensions, a strong filter and a moderate reproducer. Its four largest effects are all prohibitions: fragments, no bullets, no exclamation points, no rule of three. The measures that carry a person’s own habits — em-dash rate, paragraph asymmetry, how often the reader is addressed — move by eight to eleven points. Five measures were already at 91–97% with no file at all, and the file gets no credit for those.
<you> — three exact adjectives. The non-negotiables, first.signature phrases marked exempt.AI writing is average by design.
A model is trained to please the widest possible audience, then polished by raters who reward the safe and the fluent. So it slides toward the centre of everything it has ever read. One giveaway: humans are lumpy. Our sentence lengths swing; a model settles into one comfortable length and stays there.
Your fingerprint is made of small words.
The signal that gives a writer away doesn't live in the big content words. It lives in the glue you stopped noticing in second grade — the, of, that, while, upon, by. Stable across every subject you'll ever touch, as individual as handwriting.
In 1963, Mosteller and Wallace settled the disputed Federalist Papers with exactly this. Madison wrote whilst; Hamilton wrote while. Madison reached for upon far less often. Tiny unconscious habits, enough to assign all twelve papers with confidence that still holds.
A voiceprint is the Federalist method run backwards: measure a person, then project that fingerprint forward onto a model.
Biber found dials where everyone expected boxes.
In the late eighties the linguist Douglas Biber ran a mountain of text through factor analysis and found that features cluster into continuous dimensions. The sturdiest runs from involved (contractions, present tense, “you”) to informational (dense noun phrases, the writer stepping out of frame). A model already knows what every one of these means, so you can hand it your coordinates.
Personality leaves marks too. Lexical diversity, the ratio of adverbs to adjectives, how much you hedge. The Big Five traits show up as measurable habits. Timbrel maps all of it across thirteen dimensions.
The tells a voiceprint hunts down.
A model has a fingerprint of its own: habits it can't stop performing. Every voice file bans dozens of them by name, on by default, even in the free tier. These six are a sample, not the full list.
And your own genuine quirks? The phrases you honestly overuse are marked exempt, so they survive the cull — see it done on three real writers.
What these studies did not show.
Five studies, and two of them came out against us. This is the whole list, in the order it matters.
- The paid file does not write a better draft
- On a single draft, the free file and the paid one were indistinguishable on two of the three models tested. On the third, GPT-OSS 120B, the paid file separated sharply (13 of 15 against 5). Nobody can know in advance which case they are in, so we do not promise it. What the paid tier adds is reach and control, and that is how it is sold.
- Whether it fades late in a session is unresolved
- The numbers drift gently downward as a session runs on — 14 of 16 at the start against 9 of 16 at the end. At this sample size that is indistinguishable from noise (p = 0.23). We will not assert a decline and we will not deny one. Starting a fresh chat is prudence, not a fix for a proven failure.
- Where the rules sit in the file does nothing
- We compared the file against a byte-identical copy with the bracket blocks relocated to the middle. Indistinguishable (p = 0.55 / 0.49 / 0.38). An earlier version of this page claimed the position was the mechanism. It was wrong and the claim is retired.
- It cannot yet pull a model’s rhythm down
- Every arm that worked asked a model to write more varied. Paired against a profile whose target sits below the model’s own default, the file did not get there: 4 of 24 with no file, 4 of 23 with one. We rebuilt the engine to fix the mechanism we thought was responsible, re-ran it, and it still did not get there (p = 0.07, wrong side). Writing a model quieter is an unsolved problem here and we are not claiming otherwise.
- Six of the thirteen dimensions have no measurement at all
- Personality, humour, emotional range, narrative instinct, contrarian streak, cultural compass. No word-counter reaches them, they are not folded into any total on this page, and a draft can pass every mechanical measure here and still not sound like you. The instrument that would settle it is a blind pairing test — two drafts, one reader who knows the writer — and it does not exist yet.
- Sample sizes are small and the models keep moving
- Fifteen to thirty-two generations per cell. Every study was pre-registered before any text existed, which is what stops a small sample becoming a story, but small is small. All of it was run on models available in 2026, and a model update can move any of these numbers.
Now go and check it.