We measured our own em dashes, then banned them anyway

The em dash has become the internet's favourite proof that a machine wrote something. It is a bad proof, and the people measuring it have said so clearly. A controlled comparison puts one widely used model at about ten and a half em dashes per thousand words against a human baseline near three. The signal is in the density. One or two prove nothing about anybody.

Before we read any of that, we measured ourselves. Six thousand seven hundred words of published writing on this site: notes, tool descriptions, the about page. It came back at 2.4 em dashes per thousand words. Below the human baseline. Our notes alone were 1.5.

We were not the first here either, and this is the part that decided everything after it. On 29 April someone on this team committed "strip remaining em-dashes site-wide" and took the file holding all our prose to exactly zero. Fifteen weeks later, when we measured, that same file had thirteen. Nobody reintroduced them on purpose. They arrived one sentence at a time, in ordinary edits, none of which felt like a decision.

A cleanup is a state. It decays. That is the argument for a check rather than a tidy, and we had the evidence sitting in our own history for four months without looking at it.

The rest of the numbers were reassuring in the same way. Mean sentence length just under thirteen words. Twenty-eight percent of sentences under eight words, which is the variation that machine prose tends not to have. Zero instances across the entire corpus of the vocabulary that gets listed as a tell: no “delve”, no “tapestry”, no “testament to” anything.

So the evidence said our punctuation was fine, and we banned the character outright anyway. That decision is the interesting part, and it is worth being honest about why, because the data did not support it.

Three reasons, in ascending order of how much they mattered.

The first is that a density rule is a budget, and a budget gets spent. A rule that permits three per thousand words invites the question of whether this one is the third. Nobody wins that argument with themselves twice. A rule that permits none is decided once and never negotiated again.

The second is that the rewrite is almost always better. This is the part we did not expect. Removing seventeen em dashes meant recasting seventeen sentences, and in most of them the dash had been holding together a thought that wanted to be two sentences, or standing in for a colon that would have been more precise about the relationship. The dash is a comfortable way to avoid deciding what the second half of your sentence is doing.

We are not going to give you a count of how many came out better, because we did not assess them one by one and a number we did not measure is the thing this whole note is against. What we can say is that we did not miss a single one.

The third reason is the real one. The same rules run in more than one repository, and one of those writes copy in six languages. A rule with an exception is a rule that has to be understood before it can be applied, and it will be applied by people, and by tools, in places its author never pictured. We watched this happen: an automatic fixer built for one of these rules corrected German and French into English because it could not see which language it was editing, and it broke three translations before anyone noticed. That fixer was following a rule with a nuance in it. The nuance was correct and the fixer could not hold it.

The rule we wrote bans the en dash as well, which is the shorter mark used between numbers. So when a colleague found seven of those in a sibling repo, separating ranges like 100 Hz to 8000 Hz, and proposed an exception for a dash sitting between two numbers, we did not take it. The usage was correct typography and the exception was defensible on its own terms. It was also a hole in two other repositories that nobody would remember existed. The ranges say "to" now and read fine.

We would rather be strict about something small and cheap than carry a rule that requires judgment at every call site.

There is a version of this note that claims our writing is clean because we care more, and the measurement proves it. That version would be dishonest in a specific way. The measurement came after years of writing this way, not before, and it told us something we would not have guessed: our tool descriptions ran 5.2 per thousand words, more than three times the rate of our notes. The register we use for marketing was already drifting toward the thing we distrust, and no amount of caring had caught it. A number did.

The right question was never whether a machine wrote something. It is whether the writing is any good and whether anyone took responsibility for it. Punctuation is a weak input to that. We are strict about it because it is cheap to be strict about, not because it is the thing that matters.

Admissions stopped trying to detect, and started changing the format

Universities are switching off their AI detectors, not upgrading them. The interesting part is what they are replacing them with, and what it asks of a seventeen-year-old.

Who supplied the judgment

The provenance conversation has moved from whether a machine wrote something to who decided. We have been shipping an answer to that question, per product, in public, for a while now. Here is what it cost us to keep it honest.

We do not know how many people use it

Privacy-first is the most crowded claim in mobile right now. Ours cost us the ability to answer the first question anyone asks about a product, and we would rather describe that cost than the feature.

The gate that passed by never running

We spent a week building checks that guard our writing and our code. Four of them reported clean while checking nothing at all. Every one was found by running something, and none by reading the code.

Why we built Tennis Tutor

A junior player gets an hour of correction a week and then practises for six. The scarce thing is not court time. It is someone watching closely enough to tell you what you actually did.

Why we built Myeiyo

Chore apps either turn kids into tiny investors or turn chores into a video game. Neither matches what actually happens in a house. We built the one that does.

The decisions that don't iterate

Most things you build are reversible. A few are not. Telling them apart is harder than it sounds, and getting it wrong is what most software regret turns out to be.

What 'honest software' means in practice

We use the phrase a lot. It is easy to say. It is harder to specify.

Why we built Vyzrly

College admissions has always been a black box. We wanted to make it a little more honest.

When AI is the wrong tool

The reflex to reach for AI on every problem is a symptom of taste failure, not technical sophistication.

Why we built Glossem

Product copy lives inside code. That is a problem for everyone who is not an engineer.

Why we built USACO Tutor

Competitive programming builds a kind of thinking that matters. We wanted to make that more accessible.

Why we built ChessWarp

Every chess app asks you to find the best move. In real games, nobody tells you there is one. That gap is where most club players are stuck, and it is what we set out to fix.

Why we built Break the Test

The SAT has seven versions in circulation. Serious students burn through them in a month. The bigger problem is that even unlimited practice would not fix the thing that actually costs them points.