Road to NHS
Theme

How this is made

Editorial policy

You are trusting this with your preparation, and eventually with what you do on a ward. Nothing here is aspirational: every step is enforced by the system rather than by anybody remembering to do it.

AI may be used to assist in drafting; nothing is published until a named clinician has reviewed it.

The floor

Nothing reaches you without a doctor

A named person, not a process

Every question is signed off by a practising doctor who did not write it, whose GMC number is attached to it and shown to you. Automated checks decide how much review a draft needs. They never decide whether review happens.

Refused by the database

A reviewer cannot approve their own writing, and one whose contract, assignment, conflict-of-interest declaration or credential check is not current cannot approve anything at all. Neither is a policy document; both are constraints.

Drafting

Where a question starts

From the map, not from us

A condition, a presentation, an exam and a difficulty band, taken from the published content map rather than from our own idea of what matters.

Some are drafted with AI

And then rewritten by a doctor. Where that happened we say so on the question, name the model, and record how much of the draft survived the editing — a question you cannot trace is a question you cannot weigh.

Before a human sees it

Eight checks set the depth of review

  • Citation

    Every clinical claim cites a guideline, and that guideline still says what we said it says.

  • Numbers

    Every dose, rate, threshold and reference range in the text.

  • High-risk areas

    Prescribing, Paediatric dosing, Obstetrics, Anaesthetics and critical care, Emergency and resuscitation, Oncology, Mental Health Act and safeguarding.

  • Contradiction

    Whether the answer contradicts the explanation, or a second option is arguably also correct.

  • Give-aways

    Whether the options give the answer away.

  • Duplication

    Whether it duplicates something already in the bank.

  • Recall

    Whether it reads like a transcribed exam question.

  • House style

    UK spelling, generic drug names, SI units.

Review depth

One doctor, or two

Two doctors have to agree

Where one careful reader is not enough, because the mistake is arithmetic or the window is minutes and neither survives a second look by accident.

  • Prescribing. A dose, a route or a frequency written wrong is copied into practice verbatim.
  • Paediatric dosing. Weight-based dosing in children is where arithmetic errors become tenfold errors.
  • Anaesthetics and critical care. The margin between a therapeutic and a lethal decision is narrowest here.
  • Emergency and resuscitation. A delay caused by a wrong answer is the harm, and the window is minutes.
  • Any dose or threshold we could not check against a reference. Nothing has verified the number, so a second person must.

One doctor, specialty-matched

Everything else — including these, which are still treated as high risk, still flagged, still on the tighter deadline, and still need a written reason to approve.

  • Obstetrics. Two patients, and a drug decision that is safe in one is not in the other.
  • Oncology. Staging, urgency thresholds and referral windows decide outcomes and are frequently revised.
  • Mental Health Act and safeguarding. These are legal thresholds. A candidate who learns them wrong applies them wrong in their first week.

A second approval that has not arrived yet never discards the first. An item holding one of two approvals is in a state we call awaiting a second reviewer: the approval given is recorded, attributed and permanent, and the record shown with every question says how many approvals it needed and how many it has.

Every question needs an approval from a doctor whose specialty matches the area it belongs to. Anything containing a number is treated as higher risk before anybody looks at whether the number is right.

When a check cannot run — an unreachable guideline URL — the draft is treated as more uncertain, not less. A check that quietly passes when it did not happen is worse than no check.

What the numeric check doesNa139 mmol/L135–145K5.9 mmol/L3.5–5.3HIGHUrea6.1 mmol/L2.5–7.8Creat88 µmol/L45–90CRP4 mg/L< 5Hb134 g/L115–165

Out-of-range carries a spine, a marker and a word as well as a colour.

What the numeric check does

What the numeric check does. 6 results for an adult woman, each printed beside the reference range it is compared against. In order: Na 139 mmol/L, inside the reference range of 135–145, K 5.9 mmol/L, which is above the reference range of 3.5–5.3 and is flagged, Urea 6.1 mmol/L, inside the reference range of 2.5–7.8, Creat 88 µmol/L, inside the reference range of 45–90, CRP 4 mg/L, inside the reference range of < 5 and Hb 134 g/L, inside the reference range of 115–165. Each flagged row carries four separate signals, so that the reading survives greyscale printing and a colour-blind reader: a filled spine down its leading edge, a triangle pointing in the direction the value is out, the word HIGH at the end of the row, and the value itself in the escalation colour. This is a schematic drawn from the values listed above, not a recording from a patient.
What a triage score obligesroutineno flagsone approvalelevateda check flaggedone approval, re-checkhighhigh-risk classtwo approvals

There is no band that means no review.

What a triage score obliges

What a triage score obliges. triage sets the depth of review. It never sets whether review happens. 3 bands, stacked from the lowest score at the top: a score of routine is no flags and requires one approval, a score of elevated is a check flagged and requires one approval, re-check and a score of high is high-risk class and requires two approvals. Each band carries a filled spine down its leading edge as well as a colour, and its range, its name and its obligation are all printed. Nothing here is carried by colour alone. Nothing is placed on the bands: this is the scale, not a reading. The grammar is borrowed from the observation chart every ward uses. The cut points are our own triage’s, in lib/triage, not any published chart’s. This is a schematic drawn from the values listed above, not a recording from a patient.
The class where an unsigned box is not a clerical matterDRUGDOSEROUTESIGNEDAmoxicillin500 mgPOEnoxaparin40 mgSCunsignedParacetamol1 gPOSalbutamol5 mgNEB

An unsigned box on a drug chart is not a clerical matter. Neither is an unreviewed question about one.

The class where an unsigned box is not a clerical matter

The class where an unsigned box is not a clerical matter. an inpatient chart, ruled the way a ward chart is ruled. A ruled prescription chart with four columns — DRUG, DOSE, ROUTE and SIGNED. The rows are Amoxicillin 500 mg PO, signed, Enoxaparin 40 mg SC, with the signature box empty, Paracetamol 1 g PO, signed and Salbutamol 5 mg NEB, signed. Enoxaparin has no signature: the box carries a dashed rule and the word unsigned instead of a mark. Nothing on a ward is given because somebody meant to sign for it. Nothing on a ward is given because somebody meant to sign for it, and nothing here is published because somebody meant to review it. The doses shown are ordinary adult doses and this is not a prescribing reference. It is a picture of a signature box. This is a schematic drawn from the values listed above, not a recording from a patient.

Originality

We do not use recalled exam questions

Every stem is original, written to the published blueprint. We do not collect, buy, accept or transcribe questions from any exam, and we do not ingest anybody else’s question bank in any format. Reproducing live exam content is prohibited by the GMC and the Federation, and a bank built on it is a bank that disappears.

Triage flags drafts that read like a transcription — a paper reference left in, the vocabulary of recall, an exam’s own house phrasing — and any hit goes to the specialty lead with an explicit step confirming the item is original.

Kept current

When a guideline changes

What a reference records

Which version of a guideline the author actually read, and a quote of the passage they relied on. Each source has a re-check interval set by how fast it moves — the BNF monthly, a college guideline every six months.

What happens when it moves

Every question standing on it is flagged, ordered by how many people have seen it. A wrong question in front of two thousand candidates is a different problem from one in front of six.

When it is wrong

Reporting, and what we do about it

You report it

There is a report control on every question. It goes to the doctor who owns that area, with a deadline that depends on what you reported — a wrong answer is two days. You can see what happened to your report, in words, including when a doctor decides the question was right after all.

And we watch it ourselves

Every night the statistics on every question are recomputed. If candidates who do well overall are getting one question wrong, it is pulled from circulation automatically — there are only three explanations for that pattern and all three are defects.

The rule that pulls an item out of circulation overnightNo item has enough attempts yetdiscriminationproportion correctquarantined

The shaded band is the automatic quarantine. Nothing is plotted yet, because nothing has been answered enough times.

The rule that pulls an item out of circulation overnight

The rule that pulls an item out of circulation overnight. drawn with nothing on it, because nothing has been answered enough times yet. Every live item is one point. The horizontal axis is proportion correct — the proportion of candidates who answer it correctly — and the vertical axis is discrimination, the correlation between getting this item right and doing well on the paper overall. The band along the bottom, below a point-biserial of 0.00, is the quarantine zone. An item in it discriminates negatively: the candidates who do best overall are the ones getting it wrong. There are only three explanations for that pattern — the key is wrong, the stem is ambiguous, or the item tests something other than what it claims — and all three are defects, so it is pulled from circulation automatically that night and put in front of the doctor who owns it. Two narrower bands stand at each end of the horizontal axis: below 0.20, where almost nobody gets the item right and the key is usually wrong, and above 0.95, where almost everybody does and the item teaches nothing. Those two rules only apply once an item has at least 30 attempts, because a proportion from a dozen attempts is noise and quarantining on noise teaches everybody to ignore quarantine. Nothing is plotted on it. No item has enough attempts yet. The rule is enforced in code and runs nightly; there is simply no item yet with enough attempts on it to compute a statistic, and a scatter of invented points would be the one thing on this page that is not true. This is a schematic drawn from the values listed above, not a recording from a patient.

The honest measure

And we check ourselves

Every quarter a random sample of live questions is re-reviewed cold, by a doctor who neither wrote nor originally approved it. How often that second opinion agrees with the first is the number we treat as the honest measure of whether any of this is working. Reviewers do not know which of their decisions will be sampled, and neither do we.

You will see this mark next to reviewed content — the calibration pulse at the start of an ECG strip, the part that says the recording was measured against a known scale. Press it to see who wrote it, who signed it off, what it was checked against, and when.

The limits

What this is not

Road to NHS is an educational service. It is not affiliated with, endorsed by, or part of the NHS, the General Medical Council, or any Royal College.

It is not a clinical decision support tool, and it must not be used as one. Every question cites and links the primary guideline rather than replacing it. When you are looking after a patient, read the guideline.

Who writes this, and how they are checked