ATTENTION, PEOPLE WHOSE UNCLE SENT THEM A LIFE-CHANGING THREAD WITH NO SOURCE!

Are you tired of not rearranging your entire life around a sentence printed over a sunset?

Have you received a message beginning “They don’t want you to know this,” ending with a shopping link, three flame emojis, and no named human being who could be asked a follow-up question?

Then congratulations.

For the next several minutes, you are eligible to receive the Five Number One Rules of Knowing Whether Advice Is Any Good, a complete appraisal system containing not one, not two, but FIVE NUMBER ONE RULES.

“Wait,” you’re screaming at the screen, “they can’t all be Number One!”

That’s exactly what Big Numbering trained you to say, while it sells Rule One as a substitute for the other four.

Advice fails in specialist ways.

A clear question can’t conjure an original source.

A named source can’t make a study transfer from corporate procurement teams to your salary negotiation.

And a good fit today can’t stop a guideline, a standard, or an entire environment changing tomorrow.

Two numbers, from two very different places, to set the stakes.

The 2015 psychology replication project re-ran 100 studies. Of 97 originals reporting a significant effect, 36 replications also reached significance.

And advice that has reached the glittering gala of formal adoption can still change: a review of 363 articles testing established practices found 146 — 40.2% — reported no benefit or active harm.

Those figures don’t mean nothing works. They measure two different things in two specified samples: direct replication, and reversal of established practice.

They do mean the thread deserves more than a respectful nod before it becomes an expensive habit.

So here they are.

Five rules.

Five Number Ones.

No substitutions. No auto-renewal. No treating a prestigious logo as a force field around a claim.


RULE #1: ASK A QUESTION WITH AN OUTCOME

Introducing VIBES-TO-VERDICT™, the revolutionary home appliance that converts “I heard cold showers make people unstoppable” into a question capable of surviving contact with a clipboard.

The first job isn’t deciding whether advice sounds wise. It’s making it answerable.

State the population.

State the specific action.

State the alternative it’s supposed to beat.

State the outcome that settles it.

That’s the Ask step, and it comes first for a reason.

“Do the hardest task first thing in the morning” has no truth value yet. For whom? Doing which task, instead of what? Measured by completion rate, output quality, income, error rate, or the number of colour-coded planners purchased before lunch?

A claim about task order among knowledge workers may be an entirely different claim from task order in shift-based customer service, where the work arrives in an externally imposed sequence and nobody consulted anybody’s chronotype.

This isn’t pedantry in a tiny bow tie. It’s the difference between testing an instruction and admiring its rhythm.

And a clear question does one more thing, quietly: it tells you what evidence would count against the advice. Leave the outcome unnamed and the advice can claim victory whenever anything pleasant happens nearby.

Advice becomes testable only when it names who acts, what changes, what it is compared with, and what outcome settles the question.

An answer without a question is just a motivational poster looking for a refrigerator.


RULE #1: ACQUIRE THE ORIGINAL SOURCE

And now, from the makers of “RESEARCH SAYS”, comes the thrilling sequel: “RESEARCH IS SOMEONE ELSE’S SCREENSHOT.”

Acquire means finding the actual thing. The named study, the author, the institution, the standard, the original document.

It does not mean locating a newsletter that cites a thread that cites a podcast that mentions a person in a lab coat.

Acquire comes before Appraise because plausibility is not evidence. Without a source, all you can appraise is the sentence’s haircut.

For research, ask three questions. Who tested this? On what population? How many times, independently?

A rule with no named study, author or institution belongs at the anecdote tier no matter how confidently it says “science.” One person retelling an experience ten thousand times is still an n of one — the retellings aren’t observations.

That distinction is the whole game. Several independent sources converging on a conclusion are stronger because they carry separate original observations. Repeated copies of a single source are not convergence. They’re an echo with good distribution.

When the source is a study, there’s a ready-made tool. The Critical Appraisal Skills Programme publishes separate checklists for systematic reviews, randomised trials, cohort studies, case-control studies and qualitative research. Each asks whether the question is clear, whether the method suits it, whether the results are precise and believable, and whether they apply to the population in front of you.

It isn’t a magical truth scanner. It’s a way to stop a citation from becoming ceremonial jewellery.

A claim without a traceable source is not a conclusion waiting to be trusted. It is a provenance problem waiting to be named.

Repetition gives a claim an echo. Independent evidence gives it a floor.


RULE #1: APPRAISE THE EVIDENCE AT ITS ACTUAL TIER

BEHOLD THE TIER-O-METER DELUXE™: no chrome dial, no placebo lightning, just the faintly radical suggestion that a single observational report and a systematic review are not wearing the same hat.

Once the source is in hand, appraise the design on its own terms.

What tier is it? Does it have a suitable sample size, a comparison group, a pre-specified outcome?

A cohort study is not a randomised trial. A randomised trial is not a systematic review of several trials. A guideline built on such a review is a further institutional step again.

None of that makes any tier immune to revision. It does stop every source being graded by the font size of its conclusion.

Now keep three measurements separate, because they get blended constantly and they mean completely different things.

Direct replication asks whether an independent re-test, with comparable methods and power, again reaches significance.

Effect-size shrinkage asks how much of the original magnitude survives when a result does hold.

Guideline reversal asks whether an established practice, directly tested later, turns out to provide no benefit or net harm.

Cousins. Not interchangeable twins in matching sweaters.

The numbers make the separation matter. In that 2015 psychology project, successful replications averaged roughly half the original effect size — splitting to about 50% in cognitive psychology and about 25% in social. Eighteen laboratory economics experiments replicated at around 61%. Twenty-one social-science experiments from Nature and Science replicated at around 62%. And in a project examining 53 high-profile cancer-biology papers, effects that did replicate averaged about 85% smaller than originally reported.

No single one of those is the truth rate of all advice. They describe specified programmes in different fields, which is precisely the lesson.

A result may replicate and still be far smaller than advertised. A result may fail direct replication outright. A formally established practice may later reverse.

That 40.2% reversal figure belongs to the last category only — not to every study somebody waves around in a podcast studio.

Ask not whether the evidence has a number. Ask what that number measured, in which sample, against which failure mode.

“A study found” is the beginning of an investigation, not the closing argument.


RULE #1: APPLY ONLY ACROSS A MATCHING DOMAIN

Welcome to UNIVERSAL-ADVICE-IN-A-CAN™ — the product insisting that a tactic tested in one room will work in every room, during every decade, for every known species of payroll arrangement.

Evidence strength and domain validity are separate scores.

A source can be beautifully designed and still be a poor match for the case in front of you. Check the original population. The incentive structure. The time period.

A productivity system built around a knowledge worker’s schedule does not automatically transfer to shift work. Your situation has to share the structural features that made the evidence mean anything.

Markets give a clean version of the problem. Diversifying across uncorrelated asset classes has serious provenance — published portfolio theory, large observational datasets, multiple market cycles. That’s a different starting point from a small self-reported productivity comparison.

And yet application still depends on whether your actual holdings retain the low correlation the original analysis assumed. Correlations aren’t fixed, and they can rise sharply in a crisis.

Stronger evidence is not a travel visa.

Consensus travels imperfectly too, and here the example is unusually clarifying. The 2017 ACC/AHA hypertension guideline lowered the diagnostic threshold from 140/90 to 130/80 mmHg. Applied to the same U.S. adult population data, estimated prevalence rose from 31.9% to 45.6%.

Nobody’s blood pressure changed that morning. The boundary did.

Then comes the part the deluxe can didn’t include. The 2018 ESC/ESH guideline — along with later guidelines in Canada, Japan and Latin America — retained 140/90 for the general population, reserving 130/80 for patients already at elevated risk.

One evidence base. Different committees. Different institutional contexts. Two live positions, simultaneously.

A guideline can be genuinely useful without deciding every individual case.

A finding travels only as far as its population, incentives, assumptions, and period travel with it.

The map is valuable. It is not a teleportation device.


RULE #1: ASSESS AGAIN WHEN THE WORLD OR THE GUIDELINE CHANGES

THIS IS THE PART THEY TRIED TO SELL AS A LIFETIME WARRANTY.

From the makers of “OFFICIAL SINCE TUESDAY” comes FOREVER-CONSENSUS 9000™ — the document that supposedly becomes physically incapable of ageing the instant a committee uploads a PDF.

Formal adoption matters. Genuinely.

It doesn’t create new underlying evidence overnight, but it changes who stands behind a statement and what institutions can do with it. A consensus document gives clinicians, engineers and accountants a citable authority. It can become a reference for insurers and regulators. It can start a citation cascade in which later committees lean on it rather than re-running the same review.

So know the mechanism, not merely the tone — because the mechanisms are wildly different from each other.

The USPSTF is a panel of sixteen primary care clinicians appointed to four-year terms, grading preventive services A through D plus I for insufficient evidence. Under the Affordable Care Act, an A or B grade makes a service coverable by most private insurance without patient cost-sharing.

That’s a statutory trigger. Not the persuasive power of a confident paragraph.

GRADE works completely differently — an informal collaboration, now used or endorsed by more than 120 organisations across over 20 countries, supplying shared vocabulary for certainty and strength of recommendation. That’s a horizontal citation cascade, not a coverage mandate.

ASTM specifications are different again: testable contracts. A material either meets a numbered test method and tolerance or it doesn’t. And an ASTM document becomes mandatory only when a regulator, building code or purchasing contract cites it by number.

Same word — “standard” — three unrelated kinds of force.

Then assess after acting, and assess again when the institutional position moves. The 2002 Women’s Health Initiative results led to new FDA warning labels and a recommendation against routine combined hormone therapy for chronic-disease prevention; U.S. prescribing fell roughly 40% within about a year. Later re-analysis added nuance, particularly for women starting near menopause.

The reversal wasn’t a timeless final word either. It was the formal position at that time, doing institutional work at high speed.

Trustworthy guideline development plans periodic review rather than leaving a recommendation open-ended. Cycles and formal status vary by institution and jurisdiction — which isn’t a reason to shrug. It’s a reason to put the next check on the calendar.

Consensus is a dated institutional commitment, not an exemption from reassessment.

The most dangerous expiration date is the one nobody thinks to look for.


BUT WAIT, THERE’S MORE!

“What if it’s career advice and no randomised trial exists?”

Same rules. Ask what outcome is promised. Acquire who actually tested it. Appraise it as an anecdote, a case series, or a cohort — whatever it honestly is. Apply it only to a matching situation. Assess after use.

“What if everybody in the industry says it?”

Same rules. “Everybody” may mean independent convergence, or it may mean one charismatic PowerPoint completing a national tour. Find the original observations.

“What if it’s from a professional body?”

Same rules. Name the issuing body, the evidence grade, the date, the review process, and the external trigger that gives it consequence. A guideline is not automatically an enforceable standard.

“What if the claim is about climate, accounting, medicine, investing, or a productivity app that describes itself as an operating system?”

Same rules.

The details change.

The architecture doesn’t.


THE FIVE, WITHOUT THE SUNSET PHOTOGRAPH

Ask a question with a population, an action, a comparator, and an outcome.

Acquire the original named source before judging its evidence.

Appraise at the actual tier, keeping replication, shrinkage and reversal distinct.

Apply only where the original population, incentives, assumptions and period match yours.

Assess after acting, and re-check when the evidence, the conditions, or the formal guidance moves.


ACT NOW, WHILE SUPPLIES OF SKEPTICISM LAST

Tonight, pick one piece of advice currently renting space in your head. A diet rule, a work rule, a money rule, the life-changing thread from Uncle Caps Lock.

Write the outcome it promises. Write the population it claims to cover, and the alternative it supposedly beats.

Then go looking for its named original source. If there isn’t one, write that down rather than filling the gap with confidence. An absent source is a finding, and it’s usually the most informative one you’ll get.

If there is a study, identify its design, its comparison group, and its outcome — or take a CASP checklist to it.

Then write what differs between its world and yours. People, incentives, period, constraints.

If an institution is involved, name it precisely. Is this a guideline, a task-force grade, a GRADE recommendation, a specification, a standards update? And what mechanism actually makes it matter — a citation cascade, a coverage mandate, a code reference, an audit requirement?

Is it current? Or is it an official-looking relic with excellent typography?

Finally, set the re-check.

The goal isn’t to become incapable of acting until the entire library personally vouches for your breakfast. The goal is to match your confidence to the evidence, the fit, and the time.

A decision can be provisional and still be good.

For the low, low price of refusing to confuse familiarity with proof, the complete Five Number One system is yours.

No subscription. No shipping. No miracle exception for claims printed in a tasteful serif.

And if you act now, we’ll include the most useful tool in the box at no additional charge:

the willingness to look again.

Operators are no longer standing by.

The operator is the person deciding what to do next.