ATTENTION, PEOPLE WHO HAVE WRITTEN “Q4 WILL BE HUGE” IN A SPREADSHEET!
PEOPLE WHOSE DECK CONTAINS A HOCKEY STICK DRAWN WITH THE SERIOUSNESS OF A TREATY! PEOPLE WHO HAVE CALLED A HOPEFUL FEELING A FORECAST BECAUSE IT HAD A QUARTER NAME ATTACHED!
Are you tired of not being wrong on the record?
Have you spent another afternoon moving a finish line from “by Friday” to “by end of month” to “once the initiative lands,” while the spreadsheet nods solemnly, because spreadsheets have never been taught to ask follow-up questions?
Then congratulations.
You qualify for the Five Number One Rules of Forecasting Without a Crystal Ball — a complete prediction system for people who would like confidence to meet a reference class before confidence acquires a budget.
“Wait,” you’re shouting at the monitor, “they can’t all be Number One!”
That’s exactly what Big Numbering wants, ideally while selling you a dashboard that converts every uncertainty into a cheerful arrow.
Forecasting failures are specialists.
A perfect base rate won’t rescue a question with no deadline.
A beautiful range won’t tell you whether its author is calibrated.
A score can’t save a process that quietly redefines success after the outcome arrives.
And an update based on a new mood isn’t an update. It’s optimism arriving through the service entrance in a fake moustache.
These are gates, not toppings. The question makes something that can settle. The reference class supplies the outside view. The probability states the uncertainty. The score checks whether your confidence earns its rent. The update admits evidence without quietly replacing the forecast.
Skip one and extra enthusiasm at another does not compensate.
So here they are.
Five rules.
Five Number Ones.
No substitutions. No “we all know it’s going to happen.” No quarterly optimism performed by a committee in matching lanyards.
RULE #1: NAME THE QUESTION AND THE DEADLINE
Introducing PROPHECY-PLUS™, the revolutionary device that makes every outcome sound inevitable by declining to specify what it is.
Will revenue rise? Will the feature ship? Will the market love it?
Yes. Probably. In the broad emotional sense that things may occur somewhere.
Don’t forecast a feeling. Forecast a resolvable event.
Name a binary outcome, or a bounded quantity. Then name the deadline at which the result settles.
“Will the project ship by the stated date?” has an observable outcome. “Will the project be a success?” has none — not until somebody first decides what success means and when it gets judged.
That’s the first protection against the magical phrase on track.
On track to what?
A feature can be built and not released. Released and not adopted. A sales target can be met under a definition that was revised three weeks into the quarter, by people who felt entirely reasonable while revising it.
Each of those is a different question, with a different reference class and a different score.
The Good Judgment Project posed roughly 100 to 150 geopolitical questions a year. That volume is the point. The tournament wasn’t asking anybody to emit a general aura of geopolitical wisdom. It posed questions that closed, which is why it could score people continuously.
So write the event and the cutoff in one declarative sentence. Then keep that original wording visible after work begins.
A forecast is a commitment to a future check. Not a permission slip for a future rewrite.
A forecast that cannot lose cannot teach anything.
The deadline isn’t an inconvenience in the forecast. It’s the part that stops the forecast escaping into the woods.
RULE #1: START WITH THE REFERENCE CLASS
NOW AVAILABLE: INSIDE-VIEW DELUXE!
Comes with one vivid case study, an entire buffet of heroic anecdotes, and an enormous dial marked BUT THIS TIME IS DIFFERENT.
Batteries included. Base rate sold separately.
Pick a reference class — past cases similar enough that their outcome frequency says something about this one. State that frequency before considering the special story.
The order isn’t decorative. Class, base rate, adjustment, probability. In that sequence.
The hard part is choosing the class, and it’s genuinely hard.
Too narrow — this negotiation, between these two parties, on this issue — and you have one previous case, or none. Too broad and you’ve erased the structure that mattered.
Reference classes are judgments. Nobody discovers them lying around.
The workable convention: start broad enough to have a sample, then add one distinguishing feature at a time. Country pair. Issue type. Era. After each addition, check two things — did the base rate move, and is the sample still big enough to mean anything?
That isn’t a machine dispensing correct answers. It’s a way to stop an evocative story from posing as a denominator.
And the outside view isn’t claiming your case is ordinary. It’s insisting that “different” has to earn its adjustment.
Case-specific evidence comes after the prior, because starting with the inside view anchors everything to the story and demotes the reference class to ceremonial garnish.
The story explains why this case feels special. The reference class asks how often special cases actually cash the cheque.
Start broad. Narrow with a reason. Don’t open with an example already wearing a victory hat.
RULE #1: STATE A RANGE, NOT A VICTORY LAP
BEHOLD PROBABLY-TRON DELUXE™, the premium verbal-hedge generator.
It produces “likely,” “very likely,” and “extremely likely” — all of which can later mean precisely whatever keeps the meeting moving.
For a quantity, state an explicit range and a deadline. For a binary event, state a number.
Not a verdict. Not a word that turns to fog the moment the outcome lands.
Ending with a number is a deliberate break from punditry. Tetlock’s earlier tournaments ran from 1984 to 2003 — 284 experts, roughly 28,000 predictions — and found that experts forced to attach numbers were often only slightly better than chance, and frequently worse than simple trend extrapolation on longer horizons.
Which is the useful finding, and it isn’t flattering to anyone. A numerical commitment doesn’t make a person wise. It makes the claim inspectable, and inspectable is what was missing.
Granularity matters too. The most accurate GJP forecasters used finer distinctions than a seven-point verbal scale. 62% and 65% might both get called “probable,” but they are not the same stated belief, and across enough closed forecasts they do not produce the same record.
And don’t confuse a probability with a victory lap. A forecast of 65% says failure is still a live outcome — it is not a motivational poster with a percentage sign bolted on.
Nor does a wide interval get to impersonate rigour just because it contains every future large enough to fit in a slide.
Confidence is a number or a range with an exposed flank, not a tone of voice.
The future awards no extra points for sounding certain in a room with frosted glass.
RULE #1: SCORE THE FORECAST
FROM THE MAKERS OF “REMEMBER THAT ONE TIME I CALLED IT?” comes HIT-COUNT-O-MATIC™ — the scoring system that celebrates every lucky directional call and politely forgets the forty declarations that dissolved into mist.
Keep a ledger. Question, deadline, probability or interval, reference class, and eventually the outcome.
For binary forecasts, use the Brier score: the mean squared difference between what you said and what happened. Lower is better. Perfect is zero.
This isn’t an accuracy percentage in a lab coat, and the difference matters.
Brier is a strictly proper scoring rule. Your expected penalty is minimised by reporting the probability you actually believe — not by rounding toward certainty to look impressive.
Count-based scoring rewards overconfident rounding. Squared error makes a decisive wrong call expensive. That’s the entire design, and it’s why it works.
Then look inside the score, because one number hides two different skills.
Reliability measures the gap between your stated probabilities and observed frequencies. Resolution measures whether your forecasts actually distinguish cases from the overall base rate. Uncertainty comes from how balanced the outcomes were and isn’t yours to control.
Being well-calibrated and being discriminating are different properties. One blob called “accuracy” cannot keep them apart.
Build the calibration curve by binning your stated probabilities — ten-point bins — and plotting each bin against what actually happened. The standard is plain: among events you called 30%, about 30% should occur.
But don’t read calibration from a handful of famous calls. That curve needs hundreds of forecasts. A bin containing three of them is noise wearing a chart.
And check whether the effort bought anything: compare your score against a forecast that simply always states the historical base rate. If the two match, your adjustments added nothing — which is worth knowing, and much cheaper to learn now.
A forecast is not a prediction until it can be embarrassed by its own ledger.
Score the forecast, not the person’s ability to deliver a confident hallway monologue.
RULE #1: UPDATE FOR NEW EVIDENCE, NOT NEW MOOD
PRESENTING DEADLINE-SHIFTER PRO™, the schedule-management suite that solves every late project by moving “done” farther away until the plan is once again technically ahead of itself.
Update when the evidence is genuinely diagnostic of the defined outcome.
Start from the prior. Identify what’s actually different. Move the estimate proportionately.
That’s an adjustment.
Changing the question, the deadline, or the success measure is not an adjustment. It’s a new forecast wearing the old one’s name tag.
Goodhart is standing nearby with a fire extinguisher. The original 1975 claim was that an observed statistical regularity collapses once pressure is placed on it for control purposes — the snappy “when a measure becomes a target it ceases to be a good measure” is Strathern’s later compression, not Goodhart’s wording.
Named laws describe tendencies, not equations. The operational lesson is narrow and useful: if the team can improve the measure by redefining the measure, the score has stopped checking the thing it was built to check.
Hofstadter supplies the companion trap — work takes longer than expected, even when you account for Hofstadter’s Law. It began as a joke about forecasts that kept being restated “within ten years” without ever converging.
The joke earns its place because a revised plan is not evidence. A new date in a stronger font is not new information.
And when the proposed fix for lateness is “add people,” remember Brooks: adding manpower to a late software project makes it later. Brooks himself called that an outrageous oversimplification, then named the mechanisms — ramp-up time, communication overhead that scales combinatorially, and hard limits on how far work can be partitioned.
The rule isn’t that staffing never helps. It’s that updating your estimate requires evidence those mechanisms changed. Not a headcount-shaped wish.
New facts update a forecast. New incentives try to repaint it.
The number is allowed to move. The question doesn’t get to wander off with it.
BUT WAIT, THERE’S MORE!
“What about a sales forecast?”
Same rules. Define the revenue event and the cutoff, find comparable periods or accounts, state the range, score the record, and change it only for diagnostic evidence.
“What about a software delivery date?”
Same rules. A late project doesn’t become early because more people appeared in the status meeting. Keep the original delivery condition visible and record what actually moved the estimate.
“What about geopolitics, hiring, demand, weather, or whether the offsite will contain a trust fall?”
Same rules.
A question that closes. An outside-view base rate. A stated uncertainty. A score afterwards. An evidence-driven update.
The details change.
The architecture doesn’t.
THE FIVE, WITHOUT THE CHEERFUL ARROW
Name a binary or bounded question, and the deadline at which it resolves.
Start with a reference class and state its base rate before you consider this case’s story.
State a range or a number. Never a verbal verdict that can be reinterpreted later.
Keep the ledger and score the closed forecasts — including against a base-rate-only benchmark.
Update for diagnostic evidence while preserving the original question, deadline and scoring rule.
DO THIS TONIGHT
Open the spreadsheet containing the phrase “will be huge.”
Pick one line. Write the event that would make it true or false. Write the deadline.
If it’s a quantity, write a bounded range. If it’s binary, write a probability.
Leave the old wording in the ledger, so that future confidence can’t quietly edit the paperwork.
Then find the reference class. Start broad enough to have a sample. List the past cases. Work out the outcome frequency. Write it down as the base rate before you write anything else.
Add one distinction at a time, and only if it moves the base rate without shrinking the sample to an anecdote in business casual.
Start the ledger tonight: question, deadline, base rate, reasoning, stated probability, outcome.
Once enough forecasts close, score them. Bin the probabilities. Compare each bin against what happened. And if you haven’t reached hundreds yet, don’t treat the resulting curve as a personality test. Keep collecting.
Finally, write your update rule in advance: what evidence would genuinely change this estimate, and what would merely make the answer you want feel warmer?
Log deadline changes and definition changes separately. They may well be necessary. They’re also new forecasts, not proof the old one was right.
For the low, low price of refusing to confuse a forecast with a pep talk, the complete Five Number One system is yours.
No crystal ball. No patented certainty. No quarterly renewal of the phrase “we have strong momentum.”
Just a question, a base rate, an exposed uncertainty, a record, and the willingness to let reality return the call.
Operators are no longer standing by.
The operator is the person whose ledger this is.