ATTENTION, PEOPLE WHO HAVE JUST TURNED A SPREADSHEET INTO A TRAFFIC-LIGHT DASHBOARD!
PEOPLE WHO DISCOVERED THAT CONDITIONAL FORMATTING MAKES A CELL LOOK LIKE A DECISION! PEOPLE PREPARING TO CALL A THREE-COLOUR COLUMN “THE SYSTEM”!
Are you tired of not being a committee with a dropdown menu?
Have you spent an afternoon converting judgment into red, amber, green, and “escalate as appropriate” — as though the word appropriate did anything at all?
Then congratulations.
You qualify for the Five Number One Rules of Turning Judgment Into a Rule Without Losing the Case — the complete compression package for a choice that must become runnable without becoming a decorative superstition.
“Wait,” you’re shouting at the dashboard, “they can’t all be Number One!”
That’s exactly what Big Numbering wants: one favourite metric, promoted until it has to perform a judgment.
But failure has departments.
Error costs choose the objective.
A no-miss constraint protects the critical case.
Thresholds make the rule runnable.
Validation tests whether it travels.
And assumptions, tolerance and stop conditions tell the shortcut when to retire.
A score can’t choose its own error cost. Fit can’t certify a new population. Margin can’t repair a rule with no purpose.
So here they are.
Five rules.
Five Number Ones.
No substitutions. No “the spreadsheet says so.” No dashboard treated as a small stained-glass window through which certainty shines.
RULE #1: PRICE THE TWO KINDS OF WRONG
Introducing ERROR-O-MATIC DELUXE™, the appliance asking one rude question before it lets anybody colour a cell green:
What happens when this is wrong in each direction?
State the decision in plain language. Flag or don’t flag. Escalate or don’t. Approve or hold.
Then name both errors.
A false negative misses a case that was really positive. A false positive flags one that wasn’t.
The rule is not allowed to choose between those by accident, which is what happens whenever nobody writes them down.
Sensitivity is the share of true positives you catch. Specificity is the share of true negatives you clear.
They come as a pair because either one can be made to look heroic by ruining the other.
Flag every single case and sensitivity hits 100%, while specificity collapses to nothing. Clear every case and you get the same stunt in reverse.
Which is why a lone “accuracy” headline is a magician asking you not to inspect the other hand.
The Ottawa Ankle Rules show the choice made deliberately. Their builders started from 32 candidate findings, screened for inter-rater agreement, then partitioned under a hard constraint: 100% sensitivity for clinically significant fracture.
Not the prettiest tree. The one that accepted specificity loss rather than miss the error judged less tolerable in that setting.
And the validation numbers show what that cost. 100% sensitivity with 49% specificity in the malleolar zone. 100% sensitivity with 79% specificity in the midfoot.
Same sensitivity target. Wildly different false-alarm burden.
The pair tells a truth the first number alone would have hidden completely.
The threshold is not a fact about the world. It is a decision about which mistake gets the smaller leash.
Every traffic light contains an error budget, even when nobody has written it down.
RULE #1: PROTECT THE NO-MISS REQUIREMENT BEFORE OPTIMISING THE REST
NOW AVAILABLE: NOT-ON-MY-WATCH™ — the rule builder that refuses to trade away the case that must be caught, merely because a chart would prefer fewer alerts.
Once you’ve named the expensive error, preserve its constraint while building the shortcut.
In the Ottawa derivation, the sensitivity requirement came before the recursion picked its final cues.
That order is the whole trick. A model will happily discover a clean-looking split that improves an averaged score while quietly permitting exactly the miss the rule existed to prevent — and the averaged score will look better afterwards.
The same discipline shapes a fast-and-frugal tree.
Start with variables available at the point of decision, reduced to binary or thresholded cues. Measure each cue’s validity separately. Order them so the strongest discriminator asks first.
At each early node, one answer exits immediately and the other continues. The final cue gets two exits.
That structure is frugal for a reason, not because boxes and arrows are having a fashionable moment. A tree with m cues resolves any case in at most m questions.
Green and Mehr’s coronary-care admission tree used three cues in sequence. It matched or beat a full logistic model on that decision, while asking fewer questions on average.
Three questions. Beating the model.
Then prune the vanity branches. A node classifying only a tiny share of cases makes a slide deck feel sophisticated while contributing almost no usable coverage — fold those rare cases into the nearest surviving branch instead of adding a question that fires for four people a year.
A rule is not simple because it asks less. It is simple because every question earns its place without selling out the critical miss.
The shortest path is only clever if it still reaches the destination.
RULE #1: MAKE THE CUES FEW AND THE THRESHOLDS EXPLICIT
BEHOLD COEFFICIENT-TO-COASTER 3000™: feed in a formula only its author can operate, receive a rule that survives a meeting, a shift change, and a printer low on toner.
A rule requiring “exercise professional judgment” at every branch hasn’t compressed anything.
It has hidden the original decision inside a sentence and charged it rent.
Every cue needs an observable definition. Every continuous cue needs a cutoff. Every exit needs a stated action.
For continuous cues, automated construction typically starts at the sample median and tests a grid of candidate thresholds. That’s a construction device — not a discovered law of nature. The chosen threshold still belongs to the population, the outcome, and the error cost from Rule One.
Scores make a different bargain: check every cue, accumulate points, cross one final cutoff.
Apgar is the enduring example, and its design choices repay a look. Five bedside-observable signs, each scored 0, 1, or 2. Built for speed, without instruments.
Here’s the part people miss. Heart rate and respiratory effort later proved more informative than skin colour — and the equal weights stayed anyway.
Because equal integer weights are faster to calculate and remember under time pressure, and a score that gets computed correctly at 3 a.m. beats a better-weighted one that doesn’t.
The same mechanics convert a fitted model into a usable score. Drop predictors contributing too little. A rough planning guide of ten outcome events per predictor warns when the original model already had too many variables for its data.
Then round the surviving coefficients to small integers, and round the cutoff to something memorable — accepting a small known cost in derivation accuracy for a rule a human can actually run.
A threshold that cannot be stated cannot be checked. A score that cannot be run is only a model wearing a name tag.
Precision that vanishes at handoff was never operational precision.
RULE #1: MAKE THE RULE LOSE ITS HOME-FIELD ADVANTAGE
From the makers of “THE TRAINING DATA LOVED IT” comes AWAY-GAME VALIDATION™ — the revolutionary concept of asking whether a rule still performs after leaving the room where it learned everybody’s favourite exception.
Fit accuracy measures the data that built the rule. Prediction accuracy measures data it has never seen.
More cues and finer cutoffs almost always improve fit, because they have more ways to imitate the derivation sample.
Fit isn’t the victory lap. It’s the audition.
The German-city comparison makes the trap small and rude. Multiple regression fit its training data at about 74.3%, barely ahead of the simpler take-the-best rule at 74.2%.
Then across ten thousand cross-validated splits, take-the-best moved ahead — 72.2% against 71.9%.
The margin isn’t the lesson. The reversal is.
“More accurate” means nothing until somebody says on what data.
So keep construction and audit apart. Use one portion to choose cues, cutoffs and exits. Lock the rule. Use the other portion only to score it.
Repeat across many random splits, so one lucky partition doesn’t get its own keynote address.
And make the comparison fair — compare the compressed rule’s held-out accuracy against the full model’s held-out accuracy, on the same test split. Anything else is a before-and-after photograph taken in different weather.
Then ask whether it travels at all. The Ottawa rules went from roughly 750 patients across two hospitals, to prospective validation, to independent external review — which found sensitivity close to 100%, median specificity around 31–32%, and a 30–40% reduction in unnecessary radiography.
Those figures belong to those populations and those thresholds. They are not a promise for an unexamined setting.
A rule that has only met its own data has not been validated. It has been introduced.
Home-field advantage is not generalisation with better branding.
RULE #1: SHIP THE ASSUMPTIONS, THE MARGIN, AND THE STOP CONDITION
AND NOW, TALISMAN-IN-A-BOX™: one number, no context, and the thrilling belief that copying the number copied the engineering behind it.
Good engineering rules of thumb are useful precisely because they declare what they’re not.
A steel beam depth around L/24 is a first-pass, service-load, deflection-driven size for a model to check. A code minimum thickness may waive an explicit deflection calculation — it does not waive the strength check.
The shortcut arrives with a job, a load condition, and a boundary. All three travel together or none of them mean anything.
Margin scales with what you know. Classic design factors run 1.25–1.5 where materials and conditions are known precisely, and 2.5–4.0 where materials are untried or brittle or the environment is uncertain.
Pressure-vessel codes tell the same story institutionally: one division uses a design factor of 3.5 on ultimate tensile strength, tightened from 4.0; another permits 2.4–3.0 in exchange for more rigorous design-by-analysis.
More knowledge can buy a leaner margin. It never buys permission to forget why the margin was there.
Tolerance is the same argument in quieter clothing. Default classes apply only where a dimension has no individual tolerance — and the same class means ±0.1 mm at one size range and ±0.3 mm at another. The number is inseparable from its class and range.
And every shortcut needs an off switch. The bolt-torque expression includes a nut factor that shifts with friction — around 0.20 dry, 0.15 with machine oil, 0.12 with PTFE. That factor has no engineering derivation and must be verified on the actual fastener and finish.
It is not a reusable spell, however often it gets copied like one.
Then there’s the unpleasant sequel. The old HVAC rule of 400–600 square feet per ton ignores insulation, window area and orientation, and airtightness.
A 2021 analysis of 75 homes found actual calculated loads averaging around 1,200 square feet per ton — meaning the shortcut would specify two to three times more cooling capacity than a modern well-insulated house needs.
And the oversized unit short-cycles, dehumidifies badly, and costs more.
Conservative is not a synonym for safe. It’s just a different way to be wrong.
A shortcut earns trust by carrying its assumptions in public and its stop condition in the same box.
The rule isn’t the wisdom. The rule is the luggage tag on the wisdom.
BUT WAIT, THERE’S MORE!
“What if the dashboard is for fraud, staffing, maintenance, credit, or an incident queue?”
Same rules. Name the decision and both error costs. Protect the miss you can’t trade. Use the fewest observable cues with explicit thresholds. Test outside the data that chose them. Write down the population, tolerance, owner and stop condition.
“What if it’s an engineering calculation rather than a dashboard?”
Same rules. The error cost may be strength, serviceability, cost or schedule; the no-miss constraint may be a code requirement; validation may be analysis, test or inspection. The assumptions still travel with the shortcut.
“What if the score is already in use and everybody knows it?”
Same rules. Familiarity doesn’t tell you which error it optimises, which population built it, or whether its external validation survived a new setting. A rule can be widely used and badly overdue for its operating manual.
The details change.
The architecture doesn’t.
THE FIVE, WITHOUT THE CONDITIONAL FORMATTING
State the decision and the cost of each error before choosing any metric.
Preserve the no-miss constraint while selecting and pruning cues.
Use few observable cues, explicit thresholds, and a score or tree people can actually run.
Report sensitivity and specificity together, then test on held-out and external data.
Carry the assumptions, tolerance, margin and stop condition with the shortcut.
ACT NOW, BEFORE THE CONDITIONAL FORMATTING ACHIEVES SENTIENCE
Tonight, take one red-amber-green dashboard, or one inherited rule nobody has questioned in years.
Write its decision in a single sentence. Underneath, write the false negative and the false positive. Then mark which costs more, and why.
Don’t choose the threshold yet. That’s the discipline — the threshold is downstream of that judgment, and doing it in the other order is how the rule ends up optimising something nobody chose.
List every cue it currently uses. Cross out the ones nobody can observe reliably at the moment of decision.
For each survivor: definition, cutoff, exit, owner.
Now separate the rows that built the rule from rows that didn’t, and score the locked rule on the ones it has never seen. Report both numbers together.
If the rule has crossed into a new team, population, machine or location, that’s an external-validation question — not a launch celebration.
Finally, staple the boundary to the rule. Its population, assumptions, tolerance, margin, review owner, and the condition that retires it.
That’s the part a traffic light can’t display, which is exactly why it needs to live somewhere more durable than a colour.
For the low, low price of refusing to confuse compression with certainty, the complete Five Number One Rule Builder is yours.
No black box. No mystical coefficient. No red cell permitted to impersonate a verdict.
And if you act now, we’ll include the one component that never fits in the spreadsheet, at no additional charge:
the judgment to know when the rule has reached its edge.
Operators are no longer standing by.
The operator is the person the rule hands the case to.