A structured, severity-scored read of the methods and statistics — anchored in claim–evidence calibration, reporting standards, and the bias frameworks (RoB 2 / ROBINS-I) journals actually use. It flags concerns for you to weigh — the decisions stay yours.
Every manuscript is appraised across seven domains and severity-scored (mild, moderate, serious, critical) in each: design and claim fit, the alignment of results with the stated conclusion, statistical appropriateness, reporting-guideline adherence (the EQUATOR family — STROBE, CONSORT, PRISMA, and peers), numerical and statistical consistency, and clinical interpretability. The first six are read by the two blind engines; a seventh — contribution and literature positioning — is populated by a separate literature layer that weighs the manuscript’s novelty and framing against retrieved prior work. The overall grade is weighted by how central a problem is to the headline claim, so a serious flaw confined to a peripheral analysis does not by itself sink the paper.
For the paper’s primary conclusion, RigorMD reads two things separately: how confidently the authors state the claim (their language), and how much support the reported results actually give it — discounted for risk of bias, inconsistency, indirectness, imprecision, and publication bias (the GRADE downgrade domains) by the method, not by a vote.
The headline you see is the gap between those two, computed deterministically — not judged by the model. A humble claim backed by limited evidence reads as supported; an over-stated claim on the same evidence reads as not supported or overly confident; a claim that runs against its own results reads as counter to the results. The same evidence can earn a different headline depending only on how strongly the authors phrased the conclusion — which is the point.
Each finding is written twice. The clinician spine is one or two plain sentences — what the study can and cannot support, and the clinical “so what” — with no jargon or effect-size arithmetic. In the same main read, a “Why this matters statistically” bridge names the bias mechanism and the concrete remedy, with a demoted “Technical details” panel for the named RoB 2 / ROBINS-I domain and GRADE rationale. Read the spine to decide; read the bridge to understand why.
Two independent LLM engines appraise the manuscript blind, their reads are reconciled, and the appraisal is repeated across several passes; only findings that recur across passes are graded, so a one-off observation does not become a verdict. This is adjudication, not a single model's opinion. A deterministic layer recomputes statistics where the reported numbers allow — a flag there is a calculation you can check — and every finding carries the manuscript quote it rests on, so you can check it against your own text. Serious and critical findings are then adversarially verified: the engine tries to refute its own finding, and anything it cannot defend is withdrawn or right-sized before the report reaches you.
The blind read has also been tested against the public record — the engine independently found the documented reason PREDIMED was retracted. See the concordance benchmark.
What is a pre-submission statistical review?
A pre-submission statistical review checks a manuscript's methods and statistics before it goes to a journal — whether the study design supports the claim, whether the reported numbers are internally consistent, and whether the analysis fits the design and endpoints. RigorMD runs one automatically: two independent engines appraise the paper and a deterministic layer recomputes the statistics the reported numbers allow, returning a severity-scored, quote-grounded report.
How does RigorMD check my statistics?
A deterministic layer recomputes values from the manuscript's own reported numbers — p-values against test statistics, percentages against their denominators, and, where the reported numbers allow, means against their sample sizes — and flags the ones that do not reconcile, showing the recomputed number so you can check it yourself. Separately, the two appraisal engines weigh whether the statistical methods fit the design — the model, power, and multiplicity — and ground each concern in a quote from your text.
Does RigorMD replace peer review or a biostatistician?
No. RigorMD flags methodological and statistical concerns for you to weigh before submission; it is not a substitute for peer review or a biostatistician's input on study design. It is a pre-submission checkpoint, not a verdict.
What does RigorMD check in a manuscript?
Seven domains: design and claim fit, results-to-conclusion alignment, statistical appropriateness, reporting-guideline adherence (STROBE, CONSORT, PRISMA, and peers), numerical consistency, clinical interpretability, and contribution against the retrieved literature. Each is severity-scored from mild to critical, and every serious or critical finding is grounded in a manuscript quote or a recomputed number.
Related reading: what a pre-submission statistical review covers · every check we run · pricing.