Pre-submission methods & statistics review for clinical research

One draft in. One auditable review out.

You wrote the paper to interpret your data, draw the clinical conclusions, and make the case that the work matters — that is what reviewers should judge, not statistical errors. RigorMD gives your draft an independent, outside read that catches the distractions first: it recomputes your reported numbers, checks each claim against its evidence and the published literature, and surfaces errors, weaknesses, and opportunities to place your work correctly.

An independent, outside read — statistical and methodological. The review before the reviewers.

Or verify your statistics free, in your browser — no upload, no account →

What a chatbot can't do:your reported p = 0.04 → recomputed p = 0.078 · discordantSee every check we run →

Not sure it fits your paper? See what we'll recompute for your study →

Recomputes your statisticsReads the whole manuscriptIndependent second read — blind, reconciled16 deterministic checksClinician + statistician readEvery flag traces to your textPositioned against the literatureNo training on your files
Domain Severity Scorecardspecimen
01Design / claim fitSerious
02Results / conclusionModerate
03Statistical methodSerious
04Reporting guidelineMild
05Numerical consistencyCritical
▸ opened — the same trace stands behind every row
In your Table 2In-hospital mortality — 4.2% vs 6.1%
In your abstract“…mortality fell to 2.4%…”
What to doYou revised the table but the abstract still carries the old figure. Reconcile them before a reviewer does.
06Clinical verdictModerate
07Contribution & positioningMild
Overall Criticalrevise before submission

Built by a practicing academic surgeon and surgical journal editor with 61 peer-reviewed publications and 100+ editorial manuscript reviews.

See a sample report →
For departments, medical schools, and research offices

Every manuscript that leaves carries your program's name.

The work is the author's, but it goes out stamped with the program's name — and it represents everyone who shares it. RigorMD helps the author and protects that reputation: one consistent methodological and statistical checkpoint for trainees, faculty, and clinical research groups before journal submission.

Authors keep their reports private. Leaders see adoption, workflow status, and aggregate usage — not findings or severity. The author gets a candid review they control; the program gets a common quality floor, an audit trail for process improvement, and fewer avoidable rejection cycles.

The cheap alternative

It isn't a chatbot pass over your paper.

The cheapest thing to do is paste your manuscript into a general-purpose AI and ask if it's any good. Here's what that misses — and a reviewer won't. RigorMD is built differently: an adjudication system for clinical research — two independent engines appraise your paper blind, their findings are reconciled, serious ones are adversarially verified, and a deterministic layer neither engine controls recomputes your numbers.

A general-purpose chatbot
RigorMD built like a reviewer
Independence
A general-purpose chatbotThe same assistant that helped you write and analyze the paper — grading its own homework
RigorMDAn independent second reader: two rival vendors' engines, run blind, reconciled, and gated by a deterministic layer neither controls
Your statistics
A general-purpose chatbotReads the p-value you typed and agrees with it
RigorMDRecomputes it — χ² from your tables, denominators, percentages, means — and flags the discordance
Consistency
A general-purpose chatbotA different answer every time you ask
RigorMDAppraised across two engines and multiple passes — only findings that recur are reported, so the result is stable
Citations
A general-purpose chatbotWill invent a plausible-looking DOI
RigorMDResolves every DOI / PMID against Crossref and PubMed
Prior work
A general-purpose chatbotTakes your ‘first’ or ‘novel’ claim at face value — it retrieves no prior work to check it against
RigorMDRetrieves related prior work from PubMed and weighs your novelty framing against what it actually found — overstated firsts and unacknowledged priors flagged
Its instinct
A general-purpose chatbotTuned to be helpful and agreeable
RigorMDBuilt to find the flaw a reviewer will — and tell you
Who can act on it
A general-purpose chatbotOne voice, pitched at no one in particular
RigorMDWritten twice — a plain clinician line over a statistician panel (named bias, RoB 2 / GRADE) — so the physician and the statistician can each act on it
What you get
A general-purpose chatbotA fluent summary
RigorMDA severity-scored report where every finding ties to your quote or a recomputed number

See what the report actually looks like →

§01 — The study lifecycle
① Pre-submission review

Catch the flaw first

A severity-scored report with deterministic statistical checks — grounded in your own quotes and numbers.

$30 · per manuscript
② Re-review

Check the revised draft

After you revise, the full review runs again — compared finding-by-finding with your first report: no longer flagged, still flagged, or new.

$15 · for manuscripts reviewed here
Start free

Recompute one number yourself, free. The paid review reads the whole manuscript.

The free calculators run the same deterministic engines the paid review uses — no upload, no account. The $30 pre-submission review is the full instrument: an independent read of the whole manuscript — 16 deterministic checks plus a clinician + statistician appraisal.

Free tool

Which statistical test

Outcome type, groups, pairing, adjustment — the estimand-first recommendation, with its assumptions and its fallback.

Open the picker →
Free tool

Covariate budget (EPV)

How many covariates can your model actually support? Enter your events and candidate predictors against the events-per-variable rule.

Open the calculator →
Free tool

Reference integrity

Do your DOIs and PMIDs actually resolve — and to the work you cited?

Check your references →

All free tools →Abstract deadline calendar →

§02 — Why RigorMD

Save months. Spend less. Submit something stronger.

The publication loop is slow, statistical review is scarce, and small methodological errors are common — and costly. RigorMD finds them before a reviewer does, so revision cycles get shorter and the decision letter holds fewer surprises.

Time
6+ months to publish

Submission to publication runs from roughly 70 to 558 days across biomedical journals — and every revision or rejection round adds more. Catch the fatal flaw before a reviewer does.

Cost
Automated statistical review

Biostatistics expertise can be scarce, and academic statistical-support units report capacity constraints. Get a structured read when access is limited.

Rigor
~1 in 6 could flip the result

In one audit of orthopaedic journals, 17% of papers had a statistical error that could change the conclusion. Our deterministic layer recomputes the numbers it can and reconciles them against the text.

Integrity
Reputation is shared

Biomedical retractions have quadrupled in 20 years. The statistical and interpretive errors we flag are a preventable share — and one weak paper can shadow an entire group.

And because this is clinical research, the deepest reason is the simplest: a flawed statistic becomes flawed care. Cleaner evidence is better medicine.

§03 — How it works
STEP 01

Upload the package

Manuscript, tables, figures, supplement, cover letter, title page. PDF and DOCX. Confirm what's included.

STEP 02

Classify & adjudicate

We classify the study design and claim type, appraise the manuscript across six domains, and a separate literature layer positions it against retrieved prior work as a seventh.

STEP 03

Recompute & verify

Deterministic checks recompute p-values, denominators, percentages, and means where the reported numbers allow. Every finding is traced to a quote.

STEP 04

Severity-scored report

A reconciled, severity-scored report — emailed when ready, stored in your private dashboard.

Full validation is not instant. Most reports return by email after processing; larger packages take longer.

§04 — What we check

Seven domains, every manuscript. The first question is always: does the design support the claim?

DOMAIN 01

Design / claim fit

Whether a causal, comparative, or equivalence claim is supported by the study design actually used.

DOMAIN 02

Results / conclusion alignment

Whether the abstract, results, and conclusions agree — without spin or reframed null findings.

DOMAIN 03

Statistical appropriateness

Models, power, missing-data handling, and multiplicity, matched to the design and endpoints.

DOMAIN 04

Reporting guideline adherence

Design-specific reporting requirements, item by item, with the gaps named.

DOMAIN 05

Numerical / statistical consistency

Whether numbers, denominators, p-values, intervals, tables, and text agree with one another.

DOMAIN 06

Clinical interpretability / verdict

What a clinician or editor can legitimately take from the manuscript as written.

DOMAIN 07

Contribution & literature positioning

Whether the manuscript's novelty and framing hold up against retrieved prior work — overstated firsts and unacknowledged priors flagged.

STROBECONSORTTRIPODPRISMAROBINS-ISTARDCARESPIRIT+ more
§05 — The deliverable

Every finding shows its work.

Every serious or critical finding is grounded in a direct quote or a recomputed number — so you can verify it, not just trust it.

Worked recomputation — one finding, every stepcheck pvalue_2x2 · p-value from a 2×2 table
Quoted — manuscript, Results
“Surgical-site infection occurred in 18 of 90 patients (20.0%) in the bundle group versus 27 of 90 (30.0%) with standard care (χ² = 4.83; p = 0.028).”
a constructed specimen — and this sentence is the check’s entire input
Rebuilt — the 2×2 that sentence implies
study armSSIno SSItotal
bundle187290
standard care276390
total45135180
Step 1 · expected counts
E = row × column ÷ N
90 × 45 ÷ 180 = 22.50
90 × 135 ÷ 180 = 67.50
Equal arms — both rows expect the same.
Step 2 · contributions
(O − E)² ÷ E, per cell
(1822.50)² ÷ 22.50 = 0.900
(7267.50)² ÷ 67.50 = 0.300
(2722.50)² ÷ 22.50 = 0.900
(6367.50)² ÷ 67.50 = 0.300
Step 3 · sum, then the p-value
χ² = 0.900 + 0.300 + 0.900 + 0.300 = 2.400
df = (2 − 1) × (2 − 1) = 1
p = P(χ² ≥ 2.400) = 0.121
Reported p = 0.028 · recomputed p = 0.121 — the significance call flips at α = 0.05Critical · Domain 05
The flip is not test-dependent: Yates-corrected χ² gives p = 0.1685 and Fisher’s exact gives p = 0.1681 — every applicable test agrees. That is the engine’s own bar: when a manuscript doesn’t state its test, a flip is graded critical only if Pearson, Yates, and Fisher all flip.
A constructed specimen — no client manuscript is quoted — but the arithmetic is the product’s. This is check pvalue_2x2 as the deterministic layer runs it, graded by the engine’s own rule. CI keeps the figure honest: a TypeScript test recomputes every step above from the four counts, and the engine’s scipy path pins the same table — on every change to this site.
Figure · Discriminationfrom the PE-risk sample report
0.000.250.500.751.000.000.250.500.751.001 − SpecificitySensitivityreverse causation
Models · AUC
0.892
Dynamic · headline
+ unplanned readmission, reoperation
0.811
Static · discharge
baseline model at discharge
0.845
After washout
early PEs (1–7 d) removed
Headline gain+0.081
Survives washout+0.034

≈ 58% of the headline gain does not survive the authors' own reverse-causation check.

Illustrative ROC from the manuscript's reported AUCs. The shaded crescent is the share of the headline accuracy gain attributable to reverse causation — the report's central finding.

Read a full sample report →See every check we run →Try a free check yourself →

§06 — Built for

Three workflows. One rigor standard.

RigorMD serves authors preparing a manuscript, journals and publishers triaging submissions, and departments standardizing pre-submission QC across research groups.

§07 — Why you can trust the report

An adjudication system. An independent read, deterministic checks, and your own words.

A · 06

A reviewer's first property is independence. RigorMD had no hand in your paper — it exists only to find what a reviewer will. Each manuscript is appraised independently by two engines and reconciled into a consensus — disagreement is surfaced, not hidden. A deterministic layer recomputes statistics where the reported numbers allow, so a flag is a calculation you can check. Findings quote the manuscript directly and state, for every serious or critical issue, the direction of bias and its clinical consequence.

Tested against the public record: the engine independently found the documented reason PREDIMED was retracted →

One review, two readers. Every finding is written twice — a plain clinician line (what the study can and cannot support) over an inline “Why this matters statistically” bridge — the bias mechanism and the remedy — with the RoB 2 / ROBINS-I domain and GRADE rationale in a demoted technical panel. And the paper's headline conclusion gets a plain calibration grade: the gap between how confidently the authors state it and how much certainty the evidence can carry. A careful study stated humbly reads well; overreach is the flag, not observational data.

Fully automated, and we show our work: no human reads your manuscript as part of the analysis (operational access is limited to logged security review). Two independent engines and recomputed statistics produce a report you can verify yourself; delivery timing varies by package complexity and service availability.

And we are explicit about the limits: the report distinguishes checked and passed from not checkable from the submitted files. Without raw data, some questions can't be answered — and we say so. The report flags for your judgment and does not promise journal acceptance.

Built by a practicing academic surgeon and surgical journal editor with 61 peer-reviewed publications and 100+ editorial manuscript reviews. It exists for a simple reason: too much of the published literature does not hold up under methodological scrutiny, and the expertise that would catch the problems before submission is scarce, gated, and unevenly distributed.

This goes deeper than a general-purpose model can. The system is purpose-built on the research-on-research literature, on statistical standards, and on the reporting guidelines journals actually enforce — and its reports are written in the register physicians read. Catch the avoidable flaws before a reviewer does, and spend the months a rejection cycle costs on what your study actually shows.

§08 — Pricing

Per report. Or licensed for your organization.

① Pre-submission review · Recommended
Commissioned methods review: $150–$400
$30
Built by a practicing academic surgeon and surgical journal editor with 61 peer-reviewed publications and 100+ editorial manuscript reviews.
The severity-scored report with deterministic checks.
  • Seven-domain severity scorecard
  • Contribution & literature positioning — your novelty claims checked against retrieved prior work
  • Deterministic statistical forensics
  • Quote-grounded findings + suggested revisions
  • PDF & web report; emailed when ready
Start a review
② Institutional / editorial
Contact us
Validation infrastructure for departments, medical schools, journals, and publishers.
  • Usage-based pilots for manuscript QC programs
  • Author-private reports with administrative audit trail
  • Editorial triage workflows for journals and publishers
Department pilotEditorial Review
Departments & institutions — standardize pre-submission QC across your group.
Journals & publishers — use Editorial Review for triage before peer review.
Department pilotEditorial Review

Launch pricing — current rates apply to any report you start now. Every report is fully automated — two independent engines plus deterministic checks, with no human review in the analysis. Delivery time varies with package complexity and service availability. Pricing details & author FAQ →

§09 — Common questions

Before you upload.

What is RigorMD?

RigorMD is an automated pre-submission review for clinical research manuscripts. You upload a manuscript package — text, tables, figures, supplement — and receive a severity-scored methodological and statistical report before you submit to a journal: two independent engines appraise the paper across multiple passes, a deterministic layer recomputes your reported statistics, cited DOIs and PMIDs are resolved against Crossref and PubMed, and every finding is traced to your own text.

How is RigorMD different from pasting my manuscript into a chatbot?

A general-purpose chatbot reads your paper once and tends to agree with it — and the assistant that helped write or analyze a paper is grading its own homework. RigorMD is an independent adjudication system with no hand in your manuscript: two engines appraise the paper blind across multiple passes and only findings that recur are reported, recomputes your statistics from the reported numbers instead of trusting them, resolves every cited DOI and PMID against Crossref and PubMed, and traces each finding to your own text.

Does RigorMD actually check my statistics?

Yes. A deterministic layer recomputes p-values, denominators, percentages, and related values from the manuscript's reported numbers and flags any that do not reconcile — the recomputed number is shown in the finding.

Will it invent references the way a chatbot can?

No. Cited DOIs and PMIDs are resolved against Crossref, doi.org, and NCBI PubMed; identifiers that do not resolve, or that resolve to a different article than the one cited, are flagged in the report.

Is the report written for a clinician or a statistician?

Both. Every finding is written twice — a plain-language clinician line over a statistician panel that names the bias mechanism and its RoB 2 / GRADE domain — so the physician and the statistician can each act on it.

Does RigorMD read my paper in the context of the medical literature?

Yes. A dedicated literature layer retrieves related prior work from PubMed and weighs your framing and novelty claims against what it actually found — overstated firsts and unacknowledged prior studies are flagged, with the retrieved records cited in the report. Your manuscript is read in context, not in isolation.

Send a stronger manuscript the first time.

Upload your package and we'll email you when the severity-scored report is ready. Delivery time varies with package complexity and service availability. Submit knowing what the toughest reviewer will see, because you've already seen it.