Methods

How RigorMD reads a manuscript

A structured, severity-scored read of a manuscript's methods and statistics, organized around four plain-English questions: do the numbers hold up, does the record you cite hold up, do the conclusions match the results, and does the paper meet the reporting rules your target journal sets. It flags concerns for you to weigh — the decisions stay yours.

§01 Methods & statistical support

Whether the numbers hold up, and whether the design supports the kind of claim being made.

These four are how we organize the read. The report itself scores them as numbered domains, so a finding you see here arrives there with a severity attached.

What that looks like. Take a p-value of 0.04 reported from a 2×2 table: RigorMD recomputes that value directly from the table's own counts, using whichever test the paper names (or all three standard tests when it doesn't say), and flags it when the recomputed number — say 0.078 — lands on the other side of the 0.05 line the printed value implied. The same layer checks a reported mean against what whole-number data could actually produce: a 1–5 rating scale answered by five people cannot average to 2.53, and that kind of arithmetic gap is caught before a reviewer catches it.

What it can't reach. It works from what the manuscript reports, not the underlying dataset or code, so it cannot re-run an analysis or catch an error a fresh look at the raw data would show. Direct recomputation also only runs where the paper prints enough of the pieces to check — a percentage missing its denominator, or a mean missing its sample size, sits outside what this pillar can verify on its own.

§02 Integrity of the external record

Whether everything the paper asserts about the outside world holds — its citations, the prior work, the novelty claim.

What that looks like. A citation can look fine and still be wrong: the DOI resolves without an error, just to a different paper than the one your reference list describes. RigorMD resolves the DOIs and PMIDs you cite against Crossref, doi.org, and PubMed and flags the ones that don't match; a very long reference list or a slow registry can leave some of them unchecked, and the report says which rather than passing them. Every citation that does resolve is also checked against PubMed's and Crossref's retraction records, so a source retracted after you cited it in good faith gets flagged with the notice and its date rather than sailing through unnoticed — the same pass also catches leftover chatbot scaffolding a drafting tool left behind (“as an AI language model”) and a placeholder bracket or TODO that survived from an earlier draft. Separately, a “first to show” or “novel” claim in your introduction is weighed against what the retrieved prior literature actually reported.

What it can't reach. It can tell you a citation doesn't resolve to the paper you meant — not whether that paper, once correctly cited, is fairly characterized in your sentence. Whether a nuance from the source got flattened or oversimplified is a reading judgment for the two appraisal engines, not a database lookup. And a retraction flag doesn't read how you used the source — it can't tell a citation relied on as support from one you kept deliberately to discuss the retraction itself; that call, and what it costs your argument, stays yours.

§03 Results and conclusions in step

Whether the conclusions match what the data actually showed — tighter where the data will not carry them, bolder where they will.

For your paper's headline conclusion, RigorMD reads two things separately: how confidently you've stated it, and how much support the results actually give it, once that support is discounted for the study's own weak points — bias the design lets in, results that don't hold up consistently across analyses, an outcome that only stands in for what's actually being claimed (clinicians will recognize this bundle of considerations as the GRADE framework, named once here). A modest claim on modest support tracks the results as reported; the same support behind a bolder claim can read as stronger than the results actually give it.

What that looks like. If your text says a treatment worked in “almost all patients” while the table it cites puts the number closer to two in three, the report quotes both your sentence and the table cell it doesn't match. A subtler version of the same problem: a result reported as significant in men (p = 0.03) but not in women (p = 0.21) often gets written up as evidence the effect differs by sex — but recomputing the interaction test on those same two subgroups can come back unremarkable (p = 0.34), meaning the subgroups were never actually shown to differ from each other.

What it can't reach. It catches a conclusion drifting from the paper's own reported numbers. It doesn't independently judge whether those underlying numbers deserved confidence in the first place — a design flaw serious enough to undercut the whole result is the methods pillar's territory, not this one's.

§04 Compliance with methodology

Whether the paper meets the reporting standards its design and its target journal require.

What that looks like. For a randomized trial, that means checking the design-specific items your study calls for — allocation concealment, blinded outcome assessment, the outcomes you pre-registered — one item at a time, with each gap named rather than folded into a single overall read. Upload your target journal's own author instructions and the same pillar goes further: it reads that journal's specific submission rules and checks your manuscript against them — a randomized trial with no registration number given, an abstract missing a section the journal requires, a required statement (funding, conflicts of interest, data availability) that's simply absent — and states what would fix each one.

What it can't reach. It checks against reporting items a design or a journal actually states — it never invents a house rule a journal's instructions don't mention. Without a named target journal or its author instructions, only the general design-specific checklist runs, not that journal's own submission requirements.

§05 How you receive findings

Every finding is written twice. The clinician spine is one or two plain sentences — what the study can and cannot support, and the clinical “so what” — with no jargon or effect-size arithmetic. In the same main read, a “Why this matters statistically” bridge names the mechanism behind the concern and the concrete remedy, with the underlying bias-assessment reasoning — the RoB 2 / ROBINS-I frameworks journals already use, named here once — held in a demoted “Technical details” panel rather than the main text. Read the spine to decide; read the bridge, and the panel beneath it, to defend that decision to a co-author or a journal's statistician.

Two independent LLM engines appraise the manuscript blind and their reads are reconciled. The appraisal is then repeated across several passes, and findings generally need to recur across passes — a one-off observation does not become a verdict. Deterministic discrepancies are protected: a recomputed number is reported directly, with no recurrence hurdle. A judgment finding is instead anchored to the manuscript sentence it rests on, so you can check it against your own text. A serious or critical cross-check-only finding — raised by a single engine, not settled by math — must then survive adversarial verification: the engine tries to refute its own finding, and anything it cannot defend is withdrawn or right-sized before the report reaches you.

This section describes the author review. The editorial evaluation journals buy shares the two blind engines, the reconciliation, and the same deterministic recomputation — but not the recurrence gate or the adversarial verify pass, and that is deliberate. An editorial evaluation returns a fixed set of dimensions, each of which has to come back with a level on it, so “drop the finding” is not a move available to it; the two reads are merged by taking the more attentive one. The verify pass was closed in July 2026, after the first comparison of both products on the same manuscript: the editorial output carried close to no hard false positives, and its observed failure mode was the opposite one — false negatives from inheriting the authors' own framing, which a refutation pass does not address. See what an editor receives, and how that screen differs from this one.

§06 Scope

Scope. RigorMD flags potential methodological and statistical concerns for the authors to review. It is not a substitute for peer review or a qualified biostatistician. See a worked example on the sample report.

§07 Common questions

What is a pre-submission statistical review?

A pre-submission statistical review checks a manuscript's methods and statistics before it goes to a journal — whether the study design supports the claim, whether the reported numbers are internally consistent, and whether the analysis fits the design and endpoints. RigorMD runs one automatically: two independent engines appraise the paper and a deterministic layer recomputes the statistics the reported numbers allow, returning a severity-scored, quote-grounded report.

How does RigorMD check my statistics?

A deterministic layer recomputes values from the manuscript's own reported numbers — p-values against test statistics, percentages against their denominators, and, where the reported numbers allow, means against their sample sizes — and flags the ones that do not reconcile, showing the recomputed number so you can check it yourself. Separately, the two appraisal engines weigh whether the statistical methods fit the design — the model, power, and multiplicity — and anchor each concern to the sentence in your text it rests on.

Does RigorMD replace peer review or a biostatistician?

No. RigorMD flags methodological and statistical concerns for you to weigh before submission; it is not a substitute for peer review or a biostatistician's input on study design. It is a pre-submission checkpoint, not a verdict.

What does RigorMD check in a manuscript?

Six domains, appraised by two independent engines: design and claim fit, results-to-conclusion alignment, statistical appropriateness, reporting-guideline adherence (STROBE, CONSORT, PRISMA, and peers), numerical consistency, and clinical interpretability. A separate literature-retrieval layer positions the manuscript's novelty and framing against retrieved prior work as a seventh domain, contribution against the retrieved literature. Each is severity-scored from mild to critical, and every serious or critical finding is grounded in a manuscript quote or a recomputed number.

Related reading: what a pre-submission statistical review covers · every check we run · pricing.