Who supplied the judgment

Every tool on this site carries a field we called the AI note. It is not a disclosure badge and it is not a compliance artifact. It is a short paragraph, published on the product page, that says exactly where a machine is involved in that product and where it is not.

We added it because a label that says AI was involved tells a reader nothing they can act on. Involved how. Deciding what. The interesting information is never the presence of the thing, it is the boundary.

Some of what those paragraphs say is unflattering to us, which is the point.

The chess trainer's note says the intelligence is in the corpus rather than in any generation step: twenty six thousand positions verified by an engine before they entered the app, graded offline, with no model writing lesson content. That is a less exciting sentence than the one we could have written, and it is the true one.

The chores app for families says there is almost no machine learning in it at all, and that this was a decision rather than an oversight. A contract two people can both see, photo proof, an approval step, an append only record of what happened. Those are problems a careful schema solves better than any model. We expect to add exactly one model later, for screening submitted photos, and we said so before building it so that the addition would be visible as a change rather than arriving quietly.

The tennis tool says the vision does the seeing and nothing else. The training plan behind it is rules based, drawn from human expertise, because what a fourteen year old should work on next is a coaching judgment and not a pattern matching problem.

The wider conversation has caught up to this framing. Provenance work has moved from watermarks and detection toward attribution, and the useful distinction people have landed on is that generated and assisted are not the same thing, and that the question worth asking is who supplied the judgment. Standards bodies and regulators are building machinery to record that. We think the machinery is good and insufficient, for a reason that has nothing to do with cryptography.

A provenance record can prove a file has not been altered since signing. It cannot tell you whether the person who signed it read it.

That gap is where our field actually lives. We know this because we fell into it. This site publishes a check that reads our own copy and blocks a deploy if the writing has drifted. It passed for a week while the drift was sitting on live pages, because its scope named one file and matched one file and reported that as coverage. The check was signed, versioned, honest about its own methodology, and looking at almost nothing. Every property a provenance system would have verified was true. The claim it implied was false.

So the discipline we ended up with is narrower than a standard and more annoying to maintain. Every rule has to be shown failing before we trust it. Reintroduce the exact defect, watch the count move, put it back. Anything that claims coverage has to be able to state what it covered, out loud, in a list you can read. A green result whose scope nobody checked is worse than no check, because it stops people looking.

None of that is provenance in the technical sense. It is closer to a lab notebook. It does not prove we were right. It proves what we examined, which is the only claim any of us can honestly make.

If you want to know whether a person supplied the judgment behind something, a signature will not tell you and a detector will lie to you. What tells you is whether the thing is specific. Whether it says what it does not do. Whether it names a number and where the number came from. Whether anyone wrote down the parts that make them look bad.

We would rather be held to that standard than to a badge, and we have tried to publish enough of our own failures that holding us to it is possible.

We measured our own em dashes, then banned them anyway

The data says density is the tell, not presence, and our prose was already under the human baseline. We adopted the stricter rule regardless. Here is the argument that beat the evidence.

Admissions stopped trying to detect, and started changing the format

Universities are switching off their AI detectors, not upgrading them. The interesting part is what they are replacing them with, and what it asks of a seventeen-year-old.

We do not know how many people use it

Privacy-first is the most crowded claim in mobile right now. Ours cost us the ability to answer the first question anyone asks about a product, and we would rather describe that cost than the feature.

The gate that passed by never running

We spent a week building checks that guard our writing and our code. Four of them reported clean while checking nothing at all. Every one was found by running something, and none by reading the code.

Why we built Tennis Tutor

A junior player gets an hour of correction a week and then practises for six. The scarce thing is not court time. It is someone watching closely enough to tell you what you actually did.

Why we built Myeiyo

Chore apps either turn kids into tiny investors or turn chores into a video game. Neither matches what actually happens in a house. We built the one that does.

The decisions that don't iterate

Most things you build are reversible. A few are not. Telling them apart is harder than it sounds, and getting it wrong is what most software regret turns out to be.

What 'honest software' means in practice

We use the phrase a lot. It is easy to say. It is harder to specify.

Why we built Vyzrly

College admissions has always been a black box. We wanted to make it a little more honest.

When AI is the wrong tool

The reflex to reach for AI on every problem is a symptom of taste failure, not technical sophistication.

Why we built Glossem

Product copy lives inside code. That is a problem for everyone who is not an engineer.

Why we built USACO Tutor

Competitive programming builds a kind of thinking that matters. We wanted to make that more accessible.

Why we built ChessWarp

Every chess app asks you to find the best move. In real games, nobody tells you there is one. That gap is where most club players are stuck, and it is what we set out to fix.

Why we built Break the Test

The SAT has seven versions in circulation. Serious students burn through them in a month. The bigger problem is that even unlimited practice would not fix the thing that actually costs them points.