Admissions stopped trying to detect, and started changing the format
For three years the assumed answer to machine-written applications was better detection. That answer has quietly lost.
Inside Higher Ed reported on 5 August 2026 that Yale, Vanderbilt, Johns Hopkins and Indiana have written policies barring or discouraging faculty from treating a detector result as sole evidence of cheating, and that at least a dozen institutions including Northwestern, Georgetown and NYU have disabled Turnitin's AI detection entirely. Yale's teaching centre gave the reason plainly: the documented false positive rates are incompatible with the burden of proof an academic integrity proceeding requires. Johns Hopkins moved detection to advisory, where a hit can begin a conversation and cannot begin a charge. Curtin switched off Turnitin's AI detection at the start of this year, while leaving ordinary text matching on, which is worth separating: the plagiarism check was never the thing in dispute.
The evidence underneath those decisions is worth stating carefully, because it is routinely repeated in a stronger form than it can support. Vanderbilt's number gets quoted as seven hundred and fifty students falsely accused. It was not. Vanderbilt took Turnitin's own claimed one percent false positive rate, applied it to the seventy-five thousand papers their students submitted in 2022, and observed that around seven hundred and fifty could have been mislabelled. A projection from the vendor's own accuracy claim, published in August 2023. That is still a serious argument, and it is a different kind of argument from a count of victims.
The strongest evidence is a study rather than a policy. Liang and colleagues at Stanford, published in Patterns in 2023, ran seven commercial detectors over ninety-one TOEFL essays written by non-native English speakers under exam conditions, and eighty-eight essays by American eighth-graders. More than sixty-one percent of the TOEFL essays were classified as machine-written. The eighth-grade essays were identified as human with near-perfect accuracy. Improving word choice in the non-native samples reduced the misclassification; simplifying the native samples increased it.
Two caveats travel with that number or it becomes the Vanderbilt problem again. It tested GPT detectors generally, not Turnitin specifically. And ninety-one essays from one source is a small sample.
Even carrying both, the shape is unmistakable. The detectors were not failing randomly. They failed hardest on writing that is grammatically careful and stylistically plain, which is what a person produces in a language they acquired deliberately rather than absorbed. The tool was not identifying machines. It was identifying a particular kind of human and calling them a machine.
What replaced it is more interesting than what it replaced. UCAS is retiring the long personal statement this cycle in favour of three short structured prompts. American institutions are moving toward shorter responses, writing done in the room, interviews where a student is asked to talk about the essay they submitted, and graded classroom work used as a comparison sample.
None of that detects anything. All of it changes what the format rewards.
A long personal statement written at home over six weeks was always a test of access as much as ability: access to someone who would read three drafts and tell you the second paragraph is doing nothing. Machines made that access universal and cheap, and in doing so exposed how much of the exercise had been about access all along. The response is not to police the tool. It is to ask for the thing the tool cannot supply.
A student who can talk about their own essay for four minutes has something a generated draft does not confer. Not because talking is harder to fake, though it is. Because the question "why did you choose that example" has an answer only if a person made a choice. The new formats are all, in different ways, asking to see the judgment rather than the output.
This is the same shift happening in every field where these tools landed, and it is worth naming plainly: the question moved from did a machine write this to who supplied the judgment behind it. Those are different questions and only the second one was ever worth asking.
We build tools in this area, so we will state our interest twice over. Break the Test trains pattern recognition for standardized tests and for the essays that surround them, and the reason it exists is a bet on exactly this. Not on beating a detector. On the belief that the skill worth having is noticing what a question is actually asking, which is the skill a short structured prompt and a four minute conversation are designed to surface.
Vyzrly, which is also ours, published a longer treatment of the same shift for families and counsellors: What replaced the AI detectors. Every figure above is cited to its primary source rather than to that piece, deliberately. If either of us corrects a number later, a reader should be able to see which of us is wrong rather than find two versions of ours quietly disagreeing.
That bet looked contrarian eighteen months ago. It looks less so now.
The uncomfortable part, which we would rather write down than have pointed out, is that a shift toward interviews and in-room writing is not automatically fairer. It moves the advantage from students with access to good editors toward students who are comfortable being watched while they think. Those are different populations, and neither distribution is just. An institution that congratulates itself for retiring a biased detector and adopting a format that rewards poise has traded one bias for another and told itself a story about integrity.
The honest version of the change is narrower. Detection was measurably harming a specific group of students and producing hundreds of false accusations per campus per year. Turning it off is a clear improvement over that. What comes next is an open question that deserves the same scrutiny, and it will not get it if everyone is relieved.
For a student reading this: nobody credible is going to catch you with a detector, and the schools that tried have mostly stopped. That is not permission. It means the thing being assessed has moved to the part you cannot outsource, and the preparation that helps is the kind that makes you better at deciding, not better at producing.
Related project
Break the Test
Practice the test you actually take.
Other notes
We measured our own em dashes, then banned them anyway
The data says density is the tell, not presence, and our prose was already under the human baseline. We adopted the stricter rule regardless. Here is the argument that beat the evidence.
Who supplied the judgment
The provenance conversation has moved from whether a machine wrote something to who decided. We have been shipping an answer to that question, per product, in public, for a while now. Here is what it cost us to keep it honest.
We do not know how many people use it
Privacy-first is the most crowded claim in mobile right now. Ours cost us the ability to answer the first question anyone asks about a product, and we would rather describe that cost than the feature.
The gate that passed by never running
We spent a week building checks that guard our writing and our code. Four of them reported clean while checking nothing at all. Every one was found by running something, and none by reading the code.
Why we built Tennis Tutor
A junior player gets an hour of correction a week and then practises for six. The scarce thing is not court time. It is someone watching closely enough to tell you what you actually did.
Why we built Myeiyo
Chore apps either turn kids into tiny investors or turn chores into a video game. Neither matches what actually happens in a house. We built the one that does.
The decisions that don't iterate
Most things you build are reversible. A few are not. Telling them apart is harder than it sounds, and getting it wrong is what most software regret turns out to be.
What 'honest software' means in practice
We use the phrase a lot. It is easy to say. It is harder to specify.
Why we built Vyzrly
College admissions has always been a black box. We wanted to make it a little more honest.
When AI is the wrong tool
The reflex to reach for AI on every problem is a symptom of taste failure, not technical sophistication.
Why we built Glossem
Product copy lives inside code. That is a problem for everyone who is not an engineer.
Why we built USACO Tutor
Competitive programming builds a kind of thinking that matters. We wanted to make that more accessible.
Why we built ChessWarp
Every chess app asks you to find the best move. In real games, nobody tells you there is one. That gap is where most club players are stuck, and it is what we set out to fix.
Why we built Break the Test
The SAT has seven versions in circulation. Serious students burn through them in a month. The bigger problem is that even unlimited practice would not fix the thing that actually costs them points.