RigorMD Editorial

How the screen works

For editors who already have AI screening and want to know what is different

§00 The problem this is shaped around

A clinical reviewer reliably catches a wrong clinical claim. What survives clinical peer review is a different class of error — a confounder pushing in a knowable direction, a family of tests with no correction, an exposure defined so that survival is a precondition for entering the group. These are causal-inference errors, and they pass precisely because the work looks careful.

So the screen is not built to summarise a manuscript. It is built to name the threats to an inference that a clinician reading on their own time would not be looking for, and to say what each one obliges the office to do next.

§01 Two engines, from two vendors, read blind

Every manuscript is appraised independently by two models from different vendors — currently Anthropic and OpenAI. The vendor split is deliberate rather than incidental: the failure this product has actually shown is an engine inheriting the authors' own framing, and a second model from the same family shares the priors that produced it.

The two reads are reconciled by taking the more attentive of the two on every dimension, and the disagreement is recorded in the report rather than smoothed away — which engines ran, whether they agreed on the route, and which dimensions diverged. A report produced when only one engine could run says so, in the document. If the second engine cannot run, the evaluation degrades and discloses it; it does not silently become a single-engine read.

§02 A scale anchored on obligation, not severity

Every dimension returns one of six levels. The scale answers a single question — what does this oblige the editorial office to do before the manuscript takes its next step? It is not a severity grade, not a quality score, and not a prediction of the decision.

Clear

Nothing here changes the next step. An absence of flags, not a presence of quality — it is not “passes”, and not a prediction of the decision.

Minor

The editor should know; nothing changes. Worth a line in the decision letter.

Watch

Carry this into review. It does not change routing on its own, but a reviewer who is not told to look would plausibly read past it.

Concern

This changes what the office asks. Peer review as ordinarily assigned would not resolve it — an author clarification, one specific analysis, or a named reviewer competence is needed.

Flag

This changes whether the manuscript proceeds as submitted: scope mismatch, an ethics or registration gap, an undisclosed conflict, or a defect the submitted data cannot repair.

Not assessable

The screen could not reach a view. It means unknown, and it obliges only naming the material that would resolve it.

The rule that does the most work: a level you cannot state is closed to you. If the screen cannot name the reviewer who should look (watch), the question to ask (concern), or the gate that is failed (flag), it must use the level below. Levels are graded on the residual after the authors' own hedging and disclosure, and a level is never a decision — flag is not “reject”, clear is not “accept”.

§03 The methodological read, enumerated

The threats below are named to the engine rather than left to its discretion. Where one applies it is identified, with the quote it rests on and the reviewer competence that would settle it. Where none applies, the screen says so.

  • A confounder the design does not address and the analysis does not adjust for — with the direction it biases the result
  • Multiplicity across secondary endpoints, timepoints or subgroups with no stated correction
  • An exposure definition that requires surviving to a landmark (immortal time)
  • An odds ratio narrated as a risk, or a relative reduction quoted where events are common
  • A missing-data mechanism whose likely direction is stateable
  • A precision or equivalence claim the sample size cannot support
  • Clustering — site, surgeon, centre — left unmodelled
  • Competing risks treated as censoring
  • Intention-to-treat and per-protocol populations conflated

§04 The arithmetic, done independently

A deterministic layer recomputes statistics from the manuscript's own printed numbers — percentages against their denominators, means against their sample sizes, p-values against the tables they came from — and reports what it found, including the checks that came back consistent. It runs independently of both appraisers, so a flag there is arithmetic rather than judgment, and you can redo it yourself.

A consistent check means the printed value is reproducible from the values beside it. It does not assess the underlying data, and the report says so where it reports the checks.

§05 What we do not do, and why

The author review repeats its appraisal across several passes, gates grading on recurrence across votes, and puts findings through an adversarial verify pass. The editorial screen does neither, deliberately — and this page exists partly because saying so was not enough: scoping /methods told a journal what does not apply and nothing about what does.

An editorial evaluation returns a fixed set of dimensions, each of which has to come back carrying a level — so “drop the finding” is not a move available to it, and a recurrence gate has nothing to gate. And a refutation pass targets false positives, which is not the failure this product has shown: on the first same-manuscript comparison of both products the editorial output carried close to no hard false positives, and its observed failure mode was the opposite one — false negatives from inheriting the authors' framing. Refutation does not address that; a second vendor's independent read does.

Every report is stamped with the engine, model, prompt and schema versions that produced it, and records what reached the engine: how much text was extracted, which tables and figures the manuscript names, and how many values the numeric layer could transcribe. Where a table is an image rather than selectable text, the report says that it cannot tell a table with nothing to flag from a table it could not read.