Automated integrity checks are dependable at pattern classes: duplicated or manipulated images, statistically impossible results, retracted or fabricated references, and missing disclosures. They cannot judge intent, novelty, or scientific merit, and they cannot replace a reviewer. That boundary is why well-designed tools present findings as signals for a person to weigh, never as verdicts.
What are automated integrity checks genuinely good at?
Automated checks perform well wherever an integrity problem leaves a stable, machine-recognizable pattern. Four classes dominate editorial screening today, and each has a concrete track record.
| Pattern class | What the machine matches | Typical signal |
|---|---|---|
| Image duplication and manipulation | Pixel-level and feature-level similarity within and across figures | A reused western blot panel, a spliced lane, a rotated or rescaled duplicate |
| Impossible statistics | Arithmetic consistency between reported values | A mean that no whole-number data could produce at the stated sample size |
| Reference integrity | Bibliographies resolved against external records | A citation to a retracted paper, a DOI that resolves to nothing |
| Disclosure gaps | Presence checks against stated requirements | A missing ethics statement, data-availability section, or conflict declaration |
These patterns are common enough to justify screening every submission. Roughly 20-35% of manuscripts screened at acceptance, after peer review and before publication, get flagged for image-related issues, reports image integrity analyst Jana Christopher in an expert interview with UKRIO. On the statistics side, the GRIM test introduced by Brown and Heathers checks whether a reported mean is mathematically possible given the sample size, and around half of the testable psychology articles they examined contained at least one inconsistent reported mean. The territory keeps expanding: generative AI can now produce scientific-looking figures that experts struggle to distinguish from genuine data images, and publishers are building automated screening precisely because manual review does not reliably catch them, as Nature reported. Machines are strong in all of these cases for the same reason: the underlying task is comparison at scale, against more images, more references, and more arithmetic than any person could hold in mind. A published list of what a platform checks should map cleanly onto pattern classes like these, because those are the claims automation can actually back.
What automation cannot see: intent, merit, and context
No automated check can tell you why an anomaly exists, whether the science matters, or what a finding means inside the logic of the study. Those three blind spots, intent, merit, and context, define the ceiling of any screening system.
Intent is the clearest case. Software can establish that two image panels are identical; it cannot establish whether a researcher assembled a figure carelessly or altered it deliberately. The distinction matters in practice, because Christopher describes the ratio of unintentional error to deliberate manipulation as difficult to know for sure in the same UKRIO interview, which is exactly why a flag has to be routed to a person rather than resolved by the tool. A workflow that treats every duplication flag as an accusation misreads the base rates and corrodes trust between editors and authors.
Merit is entirely out of reach. An algorithm can verify that references resolve and that statistics are internally consistent; it cannot tell you whether the question was worth asking, whether the design was suited to it, or whether the conclusions advance the field. Those judgments are what peer review exists to make, and no screening layer replaces them.
Context sits between the two. A duplicated image may be a properly attributed republication. An inconsistent statistic may trace to a documented rounding convention. A disclosure that looks missing may live in a supplementary file. Each of these explanations is obvious to a person reading the manuscript and invisible to a pattern matcher, which is why a flag should open a conversation rather than close one.
Why the limits dictate the design
The limits of automated checking are design requirements, not caveats to bury in documentation. A system that cannot judge intent must not speak in verdicts, and that single constraint shapes everything downstream.
Three consequences follow:
- Findings are signals. The honest output of a pattern matcher is "these two regions match" or "this mean is arithmetically inconsistent": an observation that invites examination, with the strength of the match stated plainly.
- Every signal carries its evidence, traced to its source. The matched regions, the arithmetic, the retraction notice. Evidence lets a person verify or dismiss a finding in minutes instead of re-running the analysis from scratch.
- A person decides. The workflow ends with an editor, integrity officer, or reviewer who weighs the signal against the context the machine cannot see, and who owns the outcome.
This is also how the industry is organizing itself. Major publishers collaborate through the STM Integrity Hub on shared infrastructure that screens submissions for paper-mill signals and other research-integrity concerns; the screening informs editorial decisions rather than making them. The same shape should appear in any integrity platform you evaluate: signals in, evidence attached, human judgment out. A tool that knows its limits routes every finding through the one participant able to weigh intent and context, a person.
How should you evaluate an integrity tool?
Honest evaluation starts with asking where the automation stops. A vendor who cannot answer clearly has not thought hard about their own system, and a vendor who claims their system determines misconduct is claiming something no pattern matcher can deliver. Questions worth putting to any vendor:
- Which pattern classes are covered, and is the list public? Coverage claims should be specific and inspectable, tied to named check types rather than broad promises.
- What does a finding look like? Ask for a sample. It should show the evidence and where it came from, and it should read as an observation, never as a ruling.
- How is uncertainty represented? Real screening produces ambiguous cases. A tool that forces every case into flagged-or-clean is hiding its own error bars.
- Can a person overrule any finding? Every signal should be dismissible, with the reasoning recorded, because the reviewer holds context the system does not.
- What happens to the questions automation cannot assess? Merit, novelty, and intent still require reviewers and, for suspected misconduct, a formal process with due care for the people involved.
For integrity and investigation teams the bar is higher again: a finding may need to stand up in an institutional or publisher process, which makes traceable evidence a requirement rather than a nicety. The same questions apply when integrity teams assess screening support for cases already under review. A tool that is candid about its limits is more usable, not less, because candor is what lets a person defensibly own the decision. For a worked example of a check that is decisive about arithmetic and silent about intent, see the GRIM test.
How Octym helps
Octym is built around the boundary this article describes: every suspected finding is presented as a signal with its supporting evidence, each traced to its source, and a person always decides what it means and what happens next. The checks stay inside the pattern classes automation handles well; intent, merit, and context remain with editors, integrity officers, and reviewers. The reasoning behind that design is laid out in why Octym works this way.