Case study ·

Case study: the subgroup that became standard care

A trial whose main result was negative, a subgroup that wasn't, and the twenty years in between. We gave the engine the 1990 manuscript and nothing else.

RigorMD EditorialReviewed by RigorMD's founding editor — a practicing academic surgeon and surgical journal editor with 61 peer-reviewed publications and extensive editorial-review experience. About RigorMD →

§01 A trial from 1990

In 1990 a trial asked whether a steroid helped people with acute spinal cord injuries. Overall, it didn't. The groups came out the same.

But inside the trial was a smaller group — the patients who received the drug within eight hours — and they looked better. That smaller group became the standard of care. For the next two decades, if you arrived at a trauma bay with a spinal cord injury, you got the drug.

The paper is Bracken et al., (N Engl J Med 1990;322:1405-11) ↗. It has never been retracted, and nothing here suggests it should have been. This is a methods exercise in a validation context, the way the literature routinely critiques published work; it is not a commentary on any author.

§02 What the engine said, given only the paper

We submitted the manuscript and its tables through the ordinary pipeline. No commentary, no later literature, nothing about what happened next. The report came back serious, and read the paper the way its critics eventually would.

Its central finding: the headline benefit rests on a time-defined subgroup, not on the randomized comparison. In the engine's words — “the trial's overall comparison of the three groups did not show a motor benefit; the positive motor finding appears only in the patients treated within eight hours.”

Three more findings circled the same problem from different directions. The abstract foregrounds the favorable subgroup and never tells the reader the main comparison wasn't significant. Nothing in the paper shows that treating early genuinely changed the effect rather than landing that way. And with dozens of comparisons run across outcomes, timepoints and strata, a few were going to look positive on their own.

It also noticed that the strongest clinical claim — that complete injuries should no longer be considered unresponsive — comes from a smaller group inside the eight-hour group. A subgroup within the subgroup.

§03 The critics got there first

Two years after publication, neurosurgeons wrote to the Journal of Neurosurgery ↗ making the same argument: the benefit lived in a slice of the trial, not in the trial. They were right. It still took until the 2010s for the guidelines to let go of it.

That correspondence is the point of this case, not an inconvenience to it. It is independent human confirmation that the reading is the correct one — and it means nothing here depends on hindsight the engine could have absorbed. The 1992 critics didn't have hindsight either. They read the paper closely.

The only difference is when. They read it after it was published, and after it had changed how people were treated. This is the same reading, before submission.

§04 Sizing it, and what this does not show

The report graded the paper serious and its conclusions overly confident. Compare the PREDIMED case, where the republished paper discloses its own randomization problem and the same engine returned moderate with conclusions supported. Catching a problem is half the job; sizing it against what the authors already acknowledged is the other half.

What this does not show: that the engine refutes anything. It flagged that a conclusion rested on a subgroup. Whether the drug works was settled by later trials and synthesis, not by reading this paper. Nor is one case a measure of how often the engine is right — it is one report on one manuscript, chosen because the historical record gives an independent check on the answer.

What it does show is what a close, structured read of the manuscript surfaces while the paper can still be changed. Read the full sample report to see the form that read takes.