What changed our mind.

A source earns a place here by changing a specific decision in something we shipped, not by being worth reading. Each entry says which decision, so the claim can be checked.

A reading list is a claim about taste. This is a claim about work.

So it can be checked. If a source did not change something that shipped, it is not here.

GPT detectors are biased against non-native English writers

Liang, Yuksekgonul, Mao, Wu and Zou, Stanford

Seven commercial detectors run over ninety-one TOEFL essays written under exam conditions by non-native English speakers, and eighty-eight essays by American eighth-graders. More than sixty-one percent of the TOEFL essays were classified as machine-written. The eighth-grade essays were identified as human with near-perfect accuracy.

Settled that detection was the wrong thing to build against. Break the Test trains the judgment a short structured prompt is designed to surface, rather than trying to beat or evade a classifier that fails hardest on careful writing.

Break the Test

AI Detectors Are Out, New Assessments Are In

Inside Higher Ed

Yale, Vanderbilt, Johns Hopkins and Indiana barring or discouraging reliance on a detector as sole evidence; Northwestern, Georgetown and NYU disabling Turnitin's AI detection outright. Yale's stated reason: the documented false positive rates are incompatible with the burden of proof an integrity proceeding requires.

Confirmed the format shift was institutional rather than speculative, which is why our admissions writing describes what replaced detection instead of arguing about whether detection works.

Vyzrly

Occupational noise exposure limits

NIOSH and OSHA

Two different permissible exposure calculations for workplace noise, with different exchange rates. They disagree, and a dose that is acceptable under one can be well past the limit under the other.

The sound meter reports dose against both rather than picking the friendlier one. Choosing a single standard would have produced a lower number and a less useful answer.

Youth resistance training and long-term athletic development

Pediatric sports-science consensus guidance

What is appropriate for a growing athlete at a given age and stage, which is not an adult programme with the weights reduced.

Tennis Tutor's strength and mobility work is drawn from this rather than from what a professional does, and the training plan stayed rules-based instead of learned. What a fourteen-year-old should work on next is a coaching judgment.

Tennis Tutor

Stockfish

The Stockfish developers

An open-source engine strong enough to serve as ground truth for whether a tactic exists in a position, and at what depth that judgment holds.

Every one of ChessWarp's twenty-six thousand positions is verified at depth 22 before entering the corpus. The engine supplies the ground truth and the product supplies the curriculum, which is why no model writes lesson content.

ChessWarp

The long personal statement retired in favour of three short structured prompts, applying to the current admissions cycle.

Evidence that the assessment is moving toward formats where a decision has to be visible in the answer. It is the clearest external signal that the skill worth practising is noticing what a question is asking.

Break the Test