Sample report · real, unedited output on a published paper · a methods exercise, not a commentary on its authors
RigorMD Pre-submission Validation Report

A Randomized, Controlled Trial of Methylprednisolone or Naloxone in the Treatment of Acute Spinal-Cord Injury

Pre-submission check · Methodological and statistical validation against supplied materials. Authors preparing a manuscript before journal submission.

overall · SERIOUS
generated 2026-09-06
report v1
Author's snapshotat a glance — full findings follow below
Overall read
Conclusions appear directionally consistent but overstatedSerious

Fix these before you submit

Serious

Definitive efficacy claim rests on a within-8-hour subgroup, not the primary comparison

The trial's overall comparison across all patients did not show a motor benefit; the positive result the paper is built on comes only from the subset treated within eight hours. A subgroup finding cannot by itself carry a definitive 'improves recovery' conclusion.

Must-change wording

Serious

Abstract and conclusion lead with subgroup positives without stating the overall null

The abstract's takeaway is the subgroup result; a reader is not told in the same breath that when all patients are analyzed there was no significant motor benefit. That imbalance makes the treatment look more established than it is.

Must-change wording

Serious

Many unadjusted comparisons with fragile near-0.05 p-values and no multiplicity control

The paper reports dozens of separate significance tests and the key ones sit right at the 0.03–0.05 border. With that many tests and no correction, some 'significant' results are expected by chance alone.

Must-change wording

Next action

Of 9 findings: 1 needs new analysis; 7 are wording fixes; 1 is a reporting or formatting correction.

How to read this grade: severity ranks how much each concern could bear on the stated conclusions if left unaddressed — it flags what to review first. It is not a pass/fail grade of the manuscript and not a prediction of the editorial outcome.

Conclusion calibration

Conclusions appear directionally consistent but overstated

treatment with methylprednisolone in the dose used in this study improves neurologic recovery when the medication is given in the first eight hoursThe conclusion assessed, as stated in the manuscript

Authors state this with high confidence; the reported results provide low support for it. The headline grades the gap between the two — a claim–evidence calibration.

01 Design / claim fitserious
02 Results / conclusion alignmentserious
03 Statistical appropriatenessserious
04 Reporting guideline adherencemoderate
05 Numerical / statistical consistencymild
06 Clinical interpretability / verdictmoderate
07 Contribution & Literature Positioningnone
How this headline is setTwo readings are taken separately: how firmly the paper states its main conclusion, in its own words, and how far the reported evidence carries a claim of that kind. The headline grades the distance between them.

The sentence quoted above is the exact conclusion assessed. It is lifted from the manuscript rather than summarised, so an author can see precisely which claim was read.

The level is read off a fixed table rather than judged freshly each time, which is why two reports that describe the same gap land on the same words.

Some gaps here close by narrowing the conclusion to what the results already support. Others do not: where the gap comes from the analysis itself, the analysis is what has to change.

counter
Conclusions appear to run opposite to the reported results
not supported
Conclusions appear stronger than the reported evidence supports
overly confident
Conclusions appear directionally consistent but overstated
supported
Conclusions track the results as reported
solidly supported
Conclusions closely track the results as reported

§01 Domain severity scorecard

Seven questions, graded apartSeven separate questions, each graded on its own. Every finding is filed under exactly one domain, and overlapping readings of the same problem are pooled before grading. A domain's grade is its most severe finding.

A domain with no findings is a reported result, not a gap in the review: the questions in that domain were asked and nothing was raised. That is not the same as a guarantee there is nothing there to find.

Findings are also marked central or peripheral. A peripheral finding contributes at most a moderate grade to the headline, however many of them there are — so a long tail of small points cannot inflate the overall reading.

A numeric error the engine recomputed from your own reported values is always treated as central.

Of the 9 findings below, 8 ask for a wording or reporting change and 1 calls for new analysis or rework.

Seven-domain assessmenttwo-engine appraisal
DomainSeveritySummary
01Design / claim fitSeriousThe trial's overall comparison across all patients did not show a motor benefit; the positive result the paper is built on comes only from the subset treated within eight hours. (+1 more)
02Results / conclusion alignmentSeriousThe abstract's takeaway is the subgroup result; a reader is not told in the same breath that when all patients are analyzed there was no significant motor benefit. (+1 more)
03Statistical appropriatenessSeriousThe paper reports dozens of separate significance tests and the key ones sit right at the 0.03–0.05 border. (+1 more)
04Reporting guideline adherenceModerateThe Methods define the primary endpoint as the change across all three groups, but the reader cannot tell from the text whether the 8-hour cut-point and the complete/incomplete…
05Numerical / statistical consistencyMildSeveral non-negative measures (motor scores, time from accident to admission) are shown as mean ± SD where the SD is so large the distribution is clearly skewed, and several…
06Clinical interpretability / verdictModerateEven where methylprednisolone improved neurologic scores, the authors acknowledge these gains do not clearly translate into meaningful functional recovery (feeding, standing,…
07Contribution & Literature PositioningNoneNo concerns identified in this domain.
Overall severity Serious9 findings across 6 domains

§02 Strengths

What the manuscript does well — credited, not just flagged.

  • The trial used a genuinely rigorous design: multicenter, randomized, double-blind, and placebo-controlled, with two separate infusion pumps and matched placebos so that neither the methylprednisolone nor the naloxone assignment could be unblinded.
  • Follow-up was near-complete — 97.9 percent of surviving patients were examined at six weeks and 96.5 percent at six months, with mortality data on all patients — which largely removes attrition bias as an alternative explanation.
  • Baseline characteristics across the three arms were closely balanced, with no demographic or clinical variable in Table 3 differing significantly (all P>0.05), supporting successful randomization.
  • Analyses were carried out blinded and were checked for robustness with bootstrap standard errors and with independent left- and right-side neurologic scores that gave essentially identical conclusions.
  • The authors transparently report the null naloxone results and the non-significant whole-cohort motor findings alongside the positive subgroup, rather than presenting only favorable comparisons.
  • Rigorous double-blind design maintained blinding across two dissimilar active agents by giving every patient two infusions (one active or placebo per pump), so treatment allocation could not be inferred from the regimen.

§03 Findings

What severity means hereSeverity answers one question: does this change what the paper concludes? The top two grades are hard to reach on purpose — neither is available unless the report can say which way the flaw pushes the result.

Critical means the primary conclusion does not survive an honest correction. Serious means it survives but is materially weaker than stated, or a secondary conclusion is wrong.

Moderate means a real mistake that leaves every conclusion standing. Mild is the writing rather than the science — worth fixing, changing nothing.

Mild is a working grade, not a brush-off. Cosmetic items are deliberately not promoted to moderate to make the report look busier than the manuscript warrants.

A limitation you disclosed and narrowed your claim to match grades lower than one you did not. Disclosure on its own does not, if the conclusion still reaches past it.

Why each finding quotes youEvery judgment finding carries one continuous quote copied from the manuscript — no ellipses, no stitched-together fragments. A finding whose quote does not match your text is dropped before you see it.

Each quote is matched character by character against the text extracted from your file, and a finding whose quote is not found there does not reach you. So what you read is anchored to your own sentences rather than to a paraphrase of them — as far as extraction read them. It is a check on the quotation, not on the reading built from it.

Where a finding is about something missing, the quote points to where you would expect to find it — the passage that should have carried the detail. A quotation can show you the place; it cannot on its own prove nothing appears elsewhere.

Language calibration: 4 must-change wording · 2 recommended wording · 1 precision polish. Analytic work: 1 need new analysis. Submission readiness: 1 reporting/compliance correction.

01

Definitive efficacy claim rests on a within-8-hour subgroup, not the primary comparison

Serious

Design / claim fit · authors already disclose this

The trial's overall comparison across all patients did not show a motor benefit; the positive result the paper is built on comes only from the subset treated within eight hours. A subgroup finding cannot by itself carry a definitive 'improves recovery' conclusion.

Why this matters statistically

This is selective reporting of a subgroup result (RoB2 selection of the reported result): the prespecified primary test was change in score across the three groups in all randomized patients, which was not significant for motor function; significance appears only after stratifying on an 8-hour cut-point, and the paper does not report a formal treatment-by-time interaction test that would be required to claim the effect is genuinely confined to early treatment. The honest inferential status is hypothesis-generating.

Which way this pushes the result: Selecting the one subgroup (≤8 hours) in which comparisons reached significance, while the all-patient primary analysis was non-significant for motor function, inflates the apparent treatment effect.

methylprednisolone was limited to the patients treated within eight hours of their injury, supporting the hypothesis thatText used for this assessment

Review check: upheld The paper's headline conclusion — that methylprednisolone 'improves neurologic recovery when the medication is given in the first eight hours' — does rest on a subgroup in which the all-patient primary motor comparison was non-significant, and no formal treatment-by-time interaction test is reported to justify claiming the effect is confined to early treatment; correcting for this weakens the primary conclusion to hypothesis-generating without flipping it, matching the 'serious' (stretched primary conclusion) grade. The a priori stratification softens the 'selective reporting' framing somewhat but does not remove the missing interaction test, and the clinical consequence (giving steroids to all early patients) is real and actionable.

What to change

Must-change wording The current wording makes a claim the design or results cannot support.

Reframe the conclusion to state that the overall trial was non-significant for motor function and that the early-treatment benefit is a subgroup finding requiring confirmation, and report the treatment-by-time interaction p-value that underlies the ≤8-hour claim.

Technical details
Named bias
subgroup analysis · RoB2: selection of reported result
GRADE
risk of bias
Clinical consequence
Clinicians may give high-dose methylprednisolone to all early spinal-cord-injury patients on the basis of a subgroup benefit that may be a false-positive, exposing them to steroid harms (wound infection, GI bleeding) without a securely demonstrated benefit.
02

Abstract and conclusion lead with subgroup positives without stating the overall null

Serious

Results / conclusion alignment · authors already disclose this

The abstract's takeaway is the subgroup result; a reader is not told in the same breath that when all patients are analyzed there was no significant motor benefit. That imbalance makes the treatment look more established than it is.

Why this matters statistically

This is a results-conclusion alignment problem driven by selective emphasis: the Results text itself notes 'no comparable improvements in motor function' in the full cohort, but the abstract and concluding statement present only the ≤8-hour positives. Placing the overall non-significant primary comparison next to the subgroup result in the abstract would recalibrate the reader's certainty.

Which way this pushes the result: Foregrounding the favorable ≤8-hour subgroup in the abstract and conclusion while omitting that the whole-cohort motor comparison was non-significant overstates the demonstrated efficacy.

the patients who were treated with methylprednisolone within eight hours of their injury hadText used for this assessment

Review check: upheld The finding correctly identifies that the abstract and concluding statement present only the ≤8-hour subgroup positives while the whole-cohort primary motor comparison was non-significant (the Results itself notes 'No comparable improvements in motor function were observed'). The paper's primary clinical conclusion and recommendation rest on a subgroup analysis while the overall primary comparison was null; honestly stating this materially weakens the headline claim, which is the classic serious-level stretch of a primary conclusion. Partial disclosure in the Results does not bound the reliance on the subgroup, and the abstract—the part clinicians act on—omits the overall null entirely.

What to change

Must-change wording The current wording makes a claim the design or results cannot support.

Add to the abstract results the non-significant all-patient motor comparison before the ≤8-hour subgroup result, so the conclusion is read against the overall finding.

Technical details
Named bias
selective emphasis · RoB2: selection of reported result
Clinical consequence
Readers who scan only the abstract will conclude methylprednisolone broadly improves recovery and adopt it more widely than the data support.
03

Many unadjusted comparisons with fragile near-0.05 p-values and no multiplicity control

Serious

Statistical appropriateness

The paper reports dozens of separate significance tests and the key ones sit right at the 0.03–0.05 border. With that many tests and no correction, some 'significant' results are expected by chance alone.

Why this matters statistically

No multiplicity adjustment is described for the many motor/pinprick/touch × 6-week/6-month × subgroup ANOVA comparisons; the load-bearing p-values (0.033, 0.030, 0.048, 0.034) are individually fragile, and the corresponding category odds ratios (motor OR 2.04, CI 0.81–5.12; touch OR 1.69, CI 0.72–3.94) and the motor risk difference (11.8%, CI −2.9 to 25.5) have confidence intervals that cross the null—so the significance pattern is imprecise as well as unadjusted.

Which way this pushes the result: Testing three outcomes at two time points across multiple severity subgroups and time strata without any multiplicity adjustment inflates the family-wise type I error, so several 'significant' results near p=0.05 may be chance.

The P values were determined from analysis of variance.Text used for this assessment

Review check: upheld The paper's headline conclusion (methylprednisolone within 8 hours improves recovery) rests entirely on subgroup/stratified ANOVA comparisons with p-values clustered at 0.03–0.05, while the overall three-group test was not significant; no multiplicity adjustment is described despite 3 measures × 2 time points × severity strata, and the supporting odds ratios and risk difference have CIs crossing the null. Honest correction materially weakens this stretched primary conclusion, matching a serious grade.

What to change

Must-change wording The current wording makes a claim the design or results cannot support.

State explicitly that p-values are unadjusted for multiplicity, present the primary comparisons as the confirmatory analyses and the remainder as exploratory, and temper efficacy language to reflect the borderline p-values and null-crossing intervals.

Technical details
Named bias
multiple comparisons · RoB2: selection of reported result
GRADE
imprecision
Clinical consequence
A benefit that would not survive correction for multiple testing could be treated as established, leading to routine steroid use founded on statistically fragile findings.
04

Subgroup claims lack hierarchy and interaction reporting

Serious

Statistical appropriateness

The main steroid conclusion rests on a timing subgroup and several neurologic outcomes. That is clinically plausible, but the manuscript needs to show that this was the planned primary path through the data.

Why this matters statistically

The concern is selection-of-reported-result risk, not loss of randomization itself: the manuscript reports effects across treatment arms, time strata, severity strata, outcomes, and follow-up times, but does not display a testing hierarchy, multiplicity handling, or the treatment-by-time interaction that supports the eight-hour claim; this is not classic immortal-time bias because treatment timing precedes outcome assessment.

Which way this pushes the result: Inflates the apparent certainty of the within-eight-hour methylprednisolone benefit by interpreting several timing, severity, outcome, and follow-up comparisons through nominal significance.

The primary end point was a change in neurologic function between base line and the follow-up examination. Analysis of variance was used to test the hypothesis that the change in score was not different across the three treatment groups. We summarized the results using an analysis of variance for the effects of the protocol, the time the dose was received (≤8 or >8 hours from injury), and the degree of neurologic loss (complete or incomplete).Text used for this assessment

Review check: upheld The paper's headline conclusion — methylprednisolone benefit limited to the ≤8-hour window — rests on time-stratified subgroup analyses; the quoted methods describe ANOVA across treatment, time, and severity strata but report no testing hierarchy, multiplicity adjustment, or treatment-by-time interaction test, while multiple subgroup p-values hover near 0.05. Honestly correcting for this materially weakens the certainty of the primary conclusion without reversing its direction, which fits a serious grade (primary conclusion stretched).

What to change

New analysis needed The fix requires new analysis to support the claim as written.

Report the prespecified testing hierarchy and treatment-by-time interaction, then present multiplicity-controlled or hierarchical inference for motor, pinprick, and touch outcomes at the stated primary follow-up.

Technical details
Named bias
multiplicity without hierarchy · ROBINS-I: selection of reported result · RoB2: selection of reported result
GRADE
risk of bias
Clinical consequence
A clinician could adopt high-dose methylprednisolone within eight hours as a proven standard rather than a promising but less certain treatment requiring balanced discussion of benefits and harms.
05

Primary endpoint stated cohort-wide but confirmatory results only in strata; prespecification unclear

Moderate

Reporting guideline adherence

The Methods define the primary endpoint as the change across all three groups, but the reader cannot tell from the text whether the 8-hour cut-point and the complete/incomplete stratification were fully prespecified as the confirmatory analysis or added afterward.

Why this matters statistically

A CONSORT-style account of the prespecified primary analysis versus secondary/subgroup analyses is missing, as is a participant flow diagram. The authors call the time and severity hypotheses 'a priori,' but the specific 8-hour threshold and the analytic hierarchy are not documented, leaving the reader unable to distinguish confirmatory from exploratory testing—material for interpreting every downstream claim.

Which way this pushes the result: Absent a prespecified target sample size and a randomized-to-analyzed flow, the reader cannot judge whether the study was powered for the subgroup claims or how exclusions were handled.

The primary end point was a change in neurologic function between base line and the follow-up examination.Text used for this assessment

What to change

Reporting/compliance correction The fix is a reporting, guideline, disclosure, or submission-readiness correction.

State which analyses (including the 8-hour cut-point and the severity strata) were prespecified in the protocol, label the rest exploratory, and add a participant flow account of randomized, treated, and analyzed numbers per arm.

Technical details
Named bias
incomplete reporting · RoB2: selection of reported result
Clinical consequence
Reproducibility and risk-of-bias appraisal are limited; the reader cannot verify that the subgroup denominators reflect all randomized patients.
06

Neurologic-score gains do not translate directly into functional recovery

Moderate

Clinical interpretability / verdict · authors already disclose this

Even where methylprednisolone improved neurologic scores, the authors acknowledge these gains do not clearly translate into meaningful functional recovery (feeding, standing, mobility). A clinician should read the benefit as changes on an examination scale, not guaranteed real-world improvement.

Why this matters statistically

The primary endpoint is a surrogate neurologic score rather than a patient-centered functional outcome; many patients improved on several segments without changing category, so the outcome scale may overstate clinically important benefit. The authors disclose this, so the bounded interpretation is appropriate but should stay prominent.

Which way this pushes the result: The primary endpoint is a neurologic change score, not a patient-relevant functional outcome, so a clinician cannot equate the reported gains with restored independence.

the improvements in neurologic function attributable to methylprednisolone seen in the present study cannot readily be translated into specific improvements in functional statusText used for this assessment

What to change

Recommended wording The wording is directionally defensible, but softer wording would reduce reviewer risk.

Keep the disclosed surrogate-outcome caveat prominent and avoid wording that implies the neurologic-score gains equal functional recovery.

Technical details
Named bias
surrogate outcome · ROBINS-I: measurement of outcomes · RoB2: measurement of outcome
Clinical consequence
Counseling patients about functional recovery on the basis of these score changes would overstate the demonstrated benefit.
07

Nonsignificant comparisons framed as no effect

Moderate

Results / conclusion alignment · peripheral to the main claim

The negative naloxone and late-treatment results should be described as not showing benefit, not as proving no benefit. The same caution applies to safety statements, because uncommon harms can be missed.

Why this matters statistically

This is a reporting and interpretation problem: nonsignificant superiority tests are being used like equivalence or non-inferiority evidence without margins or confidence intervals that exclude clinically important effects, so the conclusion is more certain than the data allow.

Which way this pushes the result: Biases interpretation toward concluding no neurologic benefit from naloxone or later methylprednisolone and no important safety difference, when the analyses mainly show lack of statistically significant differences.

The patients treated with naloxone, or with methylprednisolone more than eight hours after their injury, did not differ in their neurologic outcomes from those given placebo. Mortality and major morbidity were similar in all three groups.Text used for this assessment

Review check: right-sized (severity reduced) The 'absence of evidence vs evidence of absence' point is real, but the naloxone secondary conclusion ('no demonstrated benefit') is a defensible reading of nonsignificant tests and the discussion actually hedges it ('No consistent evidence of efficacy was seen'); the paper also reports the actual safety numbers and P values (e.g., wound infection 7.1% vs 3.6%, P=0.21) rather than hiding underpowering, so no stated conclusion is outright wrong or flipped—this is a genuine reporting/interpretation limitation, not a secondary conclusion that is incorrect.

What to change

Must-change wording The current wording makes a claim the design or results cannot support.

Replace no-effect wording with: “Naloxone and methylprednisolone given more than eight hours after injury did not show statistically significant improvement over placebo, and the trial was not designed to prove absence of benefit or rare harms.”

Technical details
Named bias
null results as equivalence · ROBINS-I: selection of reported result · RoB2: selection of reported result
GRADE
imprecision
Clinical consequence
A clinician could rule out naloxone or late methylprednisolone, or be falsely reassured about steroid harms, when the trial did not establish equivalence or safety margins.
08

Strong 'complete injuries also benefit' claim rests on small severity subgroups

Moderate

Design / claim fit · peripheral to the main claim · authors already disclose this

The claim that even complete injuries respond, and that prior teaching 'must be reconsidered', is based on small subgroups (roughly 45 patients per arm in the plegic-total-sensory-loss cell). That is thin ground for overturning clinical expectation.

Why this matters statistically

This is a further partition within the already-selected ≤8-hour stratum (Table 5), so cell sizes fall to n≈34–47 and, for the paretic cells, to n≈11–17. Effects in such small strata are imprecise and susceptible to the same multiplicity as the main analysis; the categorical over-reach ("must be reconsidered") is selection of reported result. The direction is defensible but the certainty is not.

Which way this pushes the result: Emphasizing the complete-injury subgroup as definitively responsive overstates certainty from cells as small as 44-45 patients.

both patients with complete injuries and those with incomplete injuries improved more after treatment with methylprednisolone than after placeboText used for this assessment

What to change

Recommended wording The wording is directionally defensible, but softer wording would reduce reviewer risk.

Replace the categorical "must be reconsidered" with a bounded statement that the complete-injury subgroup was small and the finding is exploratory, and report the per-cell n alongside the claim.

Technical details
Named bias
subgroup overinterpretation · RoB2: selection of reported result
Clinical consequence
Readers may treat complete-injury patients as proven responders when the supporting data are a subgroup within a subgroup.
09

Skewed non-negative variables summarized as mean±SD; percentages lack denominators

Mild

Numerical / statistical consistency · peripheral to the main claim

Several non-negative measures (motor scores, time from accident to admission) are shown as mean ± SD where the SD is so large the distribution is clearly skewed, and several percentages are reported without their numerator/denominator, which makes the summaries harder to interpret.

Why this matters statistically

For four mean±SD summaries the mean minus two SDs falls below zero, indicating right-skew for which median (IQR) is the faithful summary; separately, ten percentages (e.g., '95 percent of whom were treated within 14 hours') lack a stated n/N (F2). Neither changes any conclusion but both reduce transparency.

Mean expanded motor score 23.7±17.4Text used for this assessment

What to change

Statistical precision The sentence is acceptable, but could be made more statistically exact.

Report skewed non-negative variables (motor scores, time to admission) as median (IQR) and give n/N for each reported percentage.

Technical details
Named bias
distributional misrepresentation

§04 Deterministic checks

Arithmetic, not judgmentNo model reasoning runs in this section. A fixed battery of arithmetic checks runs wherever the paper reports the numbers each one needs — p-values against the tables that produced them, means against what whole-number data can actually give.

Checks that came back with nothing to report are listed too, by name. A section showing only discrepancies would leave you unable to tell a check that found nothing from a check that never ran.

A number is only raised after it is confirmed to appear in the manuscript exactly as read — the guard against a misreading on our side reaching you as a finding on yours.

Statistics extracted from the manuscript and recomputed from the numbers as printed — pure arithmetic, run independently of the two appraisal engines. A consistent check means the printed value is reproducible from the other numbers the paper prints; it does not assess the underlying data. An inconsistent check also appears as a finding above. The arithmetic on any one number is exact and repeats; which numbers this layer could pull out of the manuscript does not, so the count below is what this run reached rather than a fixed property of the paper. 41 checks recomputed · 39 consistent · 2 flagged.

Check
Mean vs SD skew screen
presentation (mild): 4 of 20 mean±SD summaries of non-negative quantities have mean − 2·SD < 0, so those distributions are likely skewed (Mean expanded motor score, Mean expanded motor score, Mean expanded motor score, Time accident to admission (hours)) — median (IQR) would represent these data better
Reported vs recomputed
screened
20
flagged
4
quantities
Mean expanded motor score, Mean expanded motor score, Mean expanded motor score, Time accident to admission (hours)
rows
(1) quantity: Mean expanded motor score · mean: 23.7 · sd: 17.4 · spread label: sd · n: 161 · implied from sem: no; (2) quantity: Mean expanded motor score · mean: 24.9 · sd: 18.2 · spread label: sd · n: 153 · implied from sem: no; (3) quantity: Mean expanded motor score · mean: 24 · sd: 19.6 · spread label: sd · n: 170 · implied from sem: no; (4) quantity: Time accident to admission (hours) · mean: 3.1 · sd: 2.6 · spread label: sd · n: — · implied from sem: no
Inconsistent — flagged in findings
Check
Percentages reported without denominators
presentation (mild): 10 percentage(s) are reported without a stated denominator (e.g. "95 percent of whom were treated within 14 hours" — Abstract) — report numerator/denominator (n/N) for each percentage
Reported vs recomputed
count
10
examples
Abstract: 95% — “95 percent of whom were treated within 14 hours” · Results: 60% — “About 60 percent of the patients had complete injuries” · Results: 92% — “placed in 92 percent of the patients” · Results: 80% — “80 percent of the patients received their study drug within” · Results: 33% — “in 33 percent of the patients given methylprednisolone” · Results: 17% — “17 percent of the placebo group” · Discussion: 6% — “overall mortality rate of 6 percent in this study” · Discussion: 97.9% — “97.9 percent underwent a neurologic examination after six weeks”
Inconsistent — flagged in findings
  • Subgroup counts vs analytic N Consistent: subgroup counts sum to 487 vs analytic N=487
  • Percentage recomputation Consistent: 64/487 = 13.1% vs reported 13.1%
  • Percentage recomputation Consistent: 53/487 = 10.9% vs reported 10.9%
  • Percentage recomputation Consistent: 46/487 = 9.4% vs reported 9.4%
  • Percentage recomputation Consistent: 37/487 = 7.6% vs reported 7.6%
  • Percentage recomputation Consistent: 31/487 = 6.4% vs reported 6.4%
  • Percentage recomputation Consistent: 23/487 = 4.7% vs reported 4.7%
  • Percentage recomputation Consistent: 11/487 = 2.3% vs reported 2.3%
  • GRIM test — mean/N consistency Consistent: mean=23.7, n=161: GRIM consistent
  • GRIM test — mean/N consistency Consistent: mean=24.9, n=153: GRIM consistent
  • GRIM test — mean/N consistency Consistent: mean=24.0, n=170: GRIM consistent
  • GRIM test — mean/N consistency Consistent: mean=53.0, n=161: GRIM consistent
  • GRIM test — mean/N consistency Consistent: mean=54.5, n=154: GRIM consistent
  • GRIM test — mean/N consistency Consistent: mean=54.4, n=169: GRIM consistent
  • GRIM test — mean/N consistency Consistent: mean=54.3, n=160: GRIM consistent
  • GRIM test — mean/N consistency Consistent: mean=56.5, n=152: GRIM consistent
  • GRIM test — mean/N consistency Consistent: mean=55.7, n=168: GRIM consistent
  • Impossible p-values Consistent: p-value formatting: no impossible 'p = 0.000' values
  • Threshold-only p-value reporting Consistent: p-value formatting: exact p-values are reported (not only thresholds)
  • Over-precise p-values Consistent: p-value formatting: p-values printed at conventional precision
  • Unlabeled dispersion (SD vs SEM) Consistent: dispersion labeling: all 20 mean±spread value(s) are labeled (SD/SEM/CI/IQR)
  • Percentage precision vs sample size Consistent: percentage precision matches the denominators (10 checked)
  • Effect estimate vs its confidence interval Consistent: OR 2.04 (Results, 6-week category improvement, motor (adjusted for initial level)) lies inside its 95% CI 0.81 to 5.12
  • Effect estimate vs its confidence interval Consistent: OR 2.93 (Results, 6-week category improvement, pinprick) lies inside its 95% CI 1.26 to 6.79
  • Effect estimate vs its confidence interval Consistent: OR 1.69 (Results, 6-week category improvement, touch) lies inside its 95% CI 0.72 to 3.94
  • Effect estimate vs its confidence interval Consistent: risk difference 11.8 (Discussion, improved motor function difference (%)) lies inside its 95% CI -2.9 to 25.5
  • Effect estimate vs its confidence interval Consistent: risk difference 16.2 (Discussion, improved pinprick sensation difference (%)) lies inside its 95% CI 1.9 to 30.5
  • Effect estimate vs its confidence interval Consistent: risk difference 17.9 (Discussion, improved touch sensation difference (%)) lies inside its 95% CI 3.6 to 32.2
  • Category percentages sum to 100 Consistent: Sex — methylprednisolone (Table 3): 2 category percentages sum to 100.0% (100% within rounding)
  • Category percentages sum to 100 Consistent: Sex — naloxone (Table 3): 2 category percentages sum to 100.0% (100% within rounding)
  • Category percentages sum to 100 Consistent: Sex — placebo (Table 3): 2 category percentages sum to 100.0% (100% within rounding)
  • Category percentages sum to 100 Consistent: Race or ethnic group — methylprednisolone (Table 3): 4 category percentages sum to 100.0% (100% within rounding)
  • Category percentages sum to 100 Consistent: Extent of injury — methylprednisolone (Table 3): 2 category percentages sum to 100.0% (100% within rounding)
  • Category percentages sum to 100 Consistent: Extent of injury — naloxone (Table 3): 2 category percentages sum to 100.0% (100% within rounding)
  • Category percentages sum to 100 Consistent: Extent of injury — placebo (Table 3): 2 category percentages sum to 100.0% (100% within rounding)
  • Category percentages sum to 100 Consistent: Glasgow coma scale — methylprednisolone (Table 3): 2 category percentages sum to 100.0% (100% within rounding)
  • Category percentages sum to 100 Consistent: Glasgow coma scale — naloxone (Table 3): 2 category percentages sum to 100.0% (100% within rounding)
  • Category percentages sum to 100 Consistent: Glasgow coma scale — placebo (Table 3): 2 category percentages sum to 100.0% (100% within rounding)
  • Category percentages sum to 100 Consistent: Body mass (surface area) — methylprednisolone (Table 3): 3 category percentages sum to 99.9% (100% within rounding)

Where a table is an image rather than selectable text, its cells cannot be transcribed and the checks that would depend on them do not run. This report does not distinguish a table with nothing to flag from a table it could not read — so a short list of checks is not, on its own, evidence that the numbers are sound.

§05 Engine disagreements

Each line is one engine’s own read of a domain, taken before the two were pooled — or a place where two of this report’s own findings contradicted each other. The scorecard above shows what survived: a concern is graded only when it recurs across at least two votes and holds up to the adversarial check, so a domain can read none there while an engine flagged it here.

  • Domain 1 — Design / claim fit: the Anthropic engine graded serious; the OpenAI cross-check found nothing to flag.

§06 Journal compliance

The New England Journal of Medicinechecked against RigorMD's journal registry, as of 2026-07-02. Compliance items reflect the journal’s formatting and submission rules; they do not affect the methodological severity grade above. 8 not met · 2 could not assess · 3 met · 1 n/a.

  • ✗ Not met: Abstract word limit — expected ≤ 250 words; observed ≈ 1736 words. Reduce to at most 250 words.
  • ✗ Not met: Structured abstract — expected Four labeled sections: Background, Methods, Results, Conclusions; observed Abstract is a running narrative with no section headings (opens 'Studies in animals indicate…' and closes 'We conclude that…'). Add the four required headings and reorganize abstract content under Background, Methods, Results, and Conclusions.
  • ✗ Not met: Conflict of interest statement — expected Explicit competing-interests disclosure; observed No conflict-of-interest statement present; only drug suppliers (Upjohn, Dupont) are named in the funding note. Add a conflict-of-interest statement declaring all authors' competing interests (or their absence).
  • ✗ Not met: Author contributions statement — expected Statement describing each author's contribution; observed A personnel roster with roles (principal investigator, coinvestigator, biostatistician, etc.) is listed, but no formal author-contributions statement. Add an author-contributions statement specifying each named author's role using a standard taxonomy (e.g., CRediT).
  • ✗ Not met: AI use disclosure — expected Statement disclosing AI/generative-tool use or its absence; observed No AI-disclosure statement present. Add an AI-disclosure statement stating whether generative AI tools were used, or that none were used.
  • ✗ Not met: Trial registration statement — expected Trial registration number/statement for the clinical trial; observed No trial-registration number or statement present in the manuscript. Add a trial-registration statement citing the registry and registration number for this randomized controlled trial.
  • ✗ Not met: Clinical trial registration, protocol, and statistical analysis plan at submission — expected Registration plus protocol and SAP provided at submission; observed No registration, protocol document, or standalone statistical analysis plan is included (statistical methods are described narratively in Methods only). Provide the trial registration, the full study protocol, and the statistical analysis plan as submission files.
  • ✗ Not met: Data sharing statement for clinical trials — expected A data sharing statement; observed No data-sharing statement present. Add a data-sharing statement describing whether and how individual participant data will be made available.
  • ? Could not assess: Body word limit — expected ≤ 2700 words; observed Not checked — the Introduction-to-Conclusions span could not be identified in the extracted text, so no body word count is reported..
  • ? Could not assess: Reference style: Vancouver, numbered, cited in order — expected Vancouver numbered references cited sequentially; observed In-text citations are numbered superscripts appearing in ascending order (1, 2, 3,4, 5,6…); full reference list is not present in the extracted text.
  • ✓ Met: Tables + figures (combined) — expected ≤ 5 combined; observed 5 distinct tables/figures referenced.
  • ✓ Met: IRB/ethics statement — expected Statement of ethics/IRB approval; observed 'Institutional review boards at each center approved the study protocol.'.
  • ✓ Met: Funding statement — expected Statement of funding source; observed 'Supported by a grant (NS-15078) from the National Institute of Neurological Disorders and Stroke.'.
  • — N/A: Blinding/anonymization — expected No anonymization required (blinding: none); observed Author names, institutions, and personnel are openly listed; journal requires no blinding.

§07 Reference identifiers

No DOI or PMID identifiers were found in the manuscript text, so there was nothing to resolve against the registries. Citations without an identifier are not checkable this way.

§08 Contribution & literature positioning

Limited search — no positioning claim assessed

The literature retrieval for this manuscript returned too little to assess its positioning, so no contribution finding was made. This is not a judgment that the work is novel — the evidence pack was simply too thin to compare against.

How a finding earns its placeThe manuscript is appraised several times over, independently, using models built by two different vendors. A judgment finding is graded only when it recurs across at least two passes. Arithmetic findings stand on their own recomputation.

Passes are pooled by substance rather than by wording, so the same objection phrased two ways counts once — and counts.

Judgment findings that bear on a conclusion then go to a separate pass whose only job is to break them. Ones that cannot be defended are set aside, and the report lists them rather than dropping them quietly. Findings already settled by recomputation skip this pass — an adversarial reader must not be able to talk away arithmetic it cannot recheck.

The engine and prompt versions that produced a report are stamped on it. Two reports are only comparable if those stamps match.

Scope Legal and scope limits

Scope. Flags methodological and statistical concerns for the authors to weigh. Not medical, legal, or regulatory advice. Limited to supplied materials. Terms | Privacy | Security

RigorMD flags potential methodological and statistical concerns for the authors to review. It is not a substitute for peer review or a qualified biostatistician.

engine=0.6.1;forensics=1.0.1;anthropic=anthropic/claude-opus-4-8+reason:high,openai=openai/gpt-5.5+reason:high;prompt=appraise-0.16+recurrence-0.4+verify-0.8+extract-0.10+compliance-0.6;votes=anthropic:5,openai:1

Get this report on your manuscript.

The same severity-scored review, on a paper whose ending nobody knows yet. Upload your package and we'll email you when the report is ready; delivery time varies with package complexity and service availability.

Tested against the public record — the PREDIMED concordance benchmark →