Most images flagged during editorial screening are not fraud. Roughly 20 to 35% of manuscripts checked at acceptance raise an image-related issue, yet acceptance is ultimately rescinded for only 1 to 8% of them, per UKRIO. Misconduct requires intent, and intent cannot be read off a figure.
How often are images flagged, and for what?
Image flagging at the editorial desk is far more common than most authors realise. Between 20 and 35% of manuscripts checked at acceptance raise an image-related query during screening, while acceptance is ultimately rescinded for only 1 to 8% of them, per the same UKRIO interview. Christopher is careful not to put a number on how often a flag reflects intent, describing the ratio of unintentional error to deliberate manipulation as difficult to know for sure.
That ratio matters because it sets the correct emotional register for the whole process. A flag is a question, not a charge. When an editorial office writes to an author about a duplicated panel, the most likely explanation by a wide margin is that someone assembling a multi-panel figure in a hurry pasted the same source image into two positions, or exported the wrong version from a folder of near-identical files.
The kinds of issues that surface repeatedly are mundane:
| What screening sees | Most common explanation |
|---|---|
| The same panel appearing twice in one figure | Copy-paste error during figure assembly |
| Regions that overlap between two images | Adjacent fields of view from the same slide |
| Splice lines in a blot | Lanes rearranged for clarity, undisclosed |
| Adjusted brightness or contrast | Processing applied to make a faint band visible |
| A panel reused from an earlier paper | Legitimate reuse, missing the citation |
None of these are innocent by definition. A spliced blot can conceal a missing control, and undisclosed processing can manufacture a result. The point is that the same visual signature is produced by both careless assembly and deliberate alteration, and the image alone cannot tell you which one you are looking at.
One newer category breaks that pattern. Generative tools can now produce scientific-looking figures that experts struggle to distinguish from genuine data images, which is why publishers and integrity specialists are building automated detection rather than relying on manual review, as Nature reported. A wholly synthetic figure has no innocent explanation of the copy-paste kind, and that is precisely why it needs to be separated from the ordinary preparation errors that also surface in screening.
Which honest mistakes look most like manipulation?
Several routine laboratory practices produce images that screening software correctly flags and that an author can fully and innocently explain. Understanding them is what separates a useful screening process from an adversarial one.
Figure assembly is the largest single source. A researcher building a six-panel figure works from a folder of files with names like ctrl_2.tif and ctrl_2b.tif, and selects the wrong one. The resulting duplicate is genuine duplication: the pixels really do match. The intent behind it is a tired evening at a deadline.
Overlapping fields of view are the second. Two micrographs taken from neighbouring regions of the same slide share real content along their boundary. An automated comparison sees a partial region match, which is exactly what it is meant to see.
Blot presentation is the third. Rearranging lanes to put comparisons side by side is a long-standing practice, and it becomes a problem only when the rearrangement is not disclosed. The splice boundary is visible to detection whether or not the author intended to hide anything.
The practical consequence is that screening output should be framed as a request for the underlying data. Asking an author for the original uncropped file resolves most of these cases in a single exchange, and it does so without anyone having made an allegation. COPE guidance exists precisely to give editors a documented route for exactly this kind of query, including a flowchart for suspected image manipulation in a published article, so that handling a flag is a documented process rather than an improvised confrontation.
What does research misconduct actually require?
Research misconduct has a formal definition, and it is considerably narrower than "something is wrong with this image". The US Office of Research Integrity defines it as fabrication, falsification, or plagiarism in proposing, performing, or reviewing research, or in reporting research results.
The qualifier that does the real work is the last one: research misconduct does not include honest error or differences of opinion, again per ORI. That condition turns on what the author knew and intended, and it cannot be read off a figure.
An automated system can establish that two regions of an image match. It cannot establish that the author knew, that the duplication changed the scientific conclusion, or that accepted practice in that field was departed from significantly. Those are judgments about a person's state of mind and about disciplinary norms, and they belong to an institution's process, not to a detector.
This is why the language of screening output matters so much. A system that reports "possible duplication between Figure 2b and Figure 4a, with the matched region highlighted" gives an editor something to act on. A system that reports "manipulation detected" has made a finding it has no standing to make, and has quietly shifted a burden of proof onto an author who may have done nothing wrong. The distinction is not politeness. It is accuracy about what the evidence supports.
Why screening output should be evidence, not a verdict
The design consequence of everything above is that image screening should produce signals with their evidence attached, and route every one of them to a person. When flags are common and mostly innocent, a system that issues verdicts will be wrong often, and each wrong verdict costs an author's reputation and an editor's trust.
Evidence-first output has a specific shape. It shows the matched regions rather than asserting a conclusion about them. It names the location in the manuscript so the editor can look at the figure themselves. It ranks findings so that limited attention goes to the ones that matter most. It keeps the record of what was checked, so that a later question about the paper has something to refer back to.
This also protects the honest author, which is the constituency most often forgotten in integrity conversations. A researcher who duplicated a panel by accident is best served by a fast, specific, unaccusatory query naming the exact panels, because they can then find the correct file and fix it. A vague accusation forces them to defend their integrity instead of correcting their figure.
For editorial offices, the workflow implication is that screening belongs before peer review, where a query costs one email, rather than after publication, where the same issue costs a correction notice. Publishers building this into their submission screening are treating image checks the way they already treat plagiarism checks: a routine, expected step that most papers pass and that catches the ones that need a conversation. That normalisation is now collective as well as individual, with major publishers collaborating through the STM Integrity Hub on shared screening infrastructure. Two related cases sit outside the ordinary error caseload: figures reused from other published papers, and figures that were generated rather than photographed.
How Octym helps
Octym reviews figures for manipulation, duplication, and reuse, and reports what it finds as signals with the evidence attached: the matched regions, their locations in the manuscript, and their severity. Image and figure integrity is powered by Proofig, the largest image-plagiarism database in the industry. Nothing Octym returns is an accusation, and no finding closes a question on its own; a person always decides what a flagged figure means. You can see the full set of checks on the what we check page.