Figure reuse across different papers is caught by comparing submitted images against the published record at scale: software extracts figures, graphs, and charts from a manuscript and matches them against a database of previously published material. Manual review cannot do this, because no editor or reviewer holds millions of published figures in memory.
Why does cross-paper figure reuse escape manual review?
Manual image screening works within the boundaries of a single manuscript. An editor or image-integrity analyst can place two figures side by side, notice a repeated blot lane or a duplicated microscopy field, and query the authors before anything is published. That checking is real and productive: image-integrity analyst Jana Christopher has estimated that roughly 20 to 35% of the accepted manuscripts she screened before publication were flagged for image-related problems, as she described in an expert interview with UKRIO.
The limit is the comparison set. Within-manuscript screening compares a paper against itself: a few dozen images, all on one desk. Reuse across papers requires comparing a submission against everything the literature has already published, and no human can do that. A reviewer reads a handful of manuscripts a year and recognizes, at best, figures from papers they happen to know well. A chart lifted from another group's article, published years earlier in a journal the reviewer never opens, looks entirely original.
That is why cross-paper reuse tends to surface by accident: a reader with an unusual visual memory, a post-publication sleuth, a tip. Screening that depends on coincidence is not screening. The scale mismatch is structural rather than a matter of effort: human attention grows with hours spent reading, while the published record already runs to millions of articles. Closing that gap is a database problem, and it has to be solved as one.
What does image plagiarism look like in practice?
Image plagiarism in practice covers more than photographs. The obvious cases are photographic data images: microscopy fields, western blots, gels, and flow-cytometry plots taken from a published paper and presented as new results, often rotated, cropped, mirrored, or contrast-adjusted along the way. The overlooked cases are graphs and charts. A bar chart or line graph is an image like any other, and a lifted plot with relabeled axes and recolored series passes text-similarity screening untouched, because text tools compare words and see nothing inside a figure.
| What gets reused | Typical alterations | Why readers miss it |
|---|---|---|
| Microscopy, blot, and gel images | Rotation, mirroring, cropping, contrast or color shifts | An altered copy does not look identical at reading speed |
| Graphs and charts | Relabeled axes, recolored series, redrawn legends | Treated as results rather than images, so no one compares them visually |
| Panels inside compound figures | One panel extracted and given a new caption | The source is a sub-image buried in a different figure in a different paper |
The problem now runs in both directions: generative AI can produce scientific-looking figures that experts struggle to distinguish from genuine data images, as Nature reported in 2024. A screen for reused or fabricated figures therefore has to treat every visual element in a manuscript as comparable data: photographs, plots, charts, and the individual panels inside compound figures alike.
How does comparison against the published record work?
Comparing a submission against the published record is a database problem with three parts. First, a corpus: an indexed collection of figures extracted from published articles, the material every new submission is compared against. Second, segmentation: compound figures are split into panels, because reuse usually involves one panel rather than a whole figure. Third, matching that survives the standard alterations: each image is reduced to a visual fingerprint that stays stable under rotation, mirroring, scaling, cropping, and contrast changes, so a recolored chart or a flipped micrograph still finds its source.
The corpus sets the ceiling. A comparison can only surface matches in material it has indexed, so the breadth of the database determines what the screen can see.
Publishers have already accepted this logic for neighboring problems: through the STM Integrity Hub, major publishers collaborate on shared infrastructure that screens submissions for paper-mill signals and other research-integrity concerns, because no single journal can see across the whole submission stream. The same pressure is reaching figures: publishers and integrity specialists are developing automated screening because manual review does not reliably catch AI-generated images, according to Nature news.
The output of a database comparison is a shortlist of candidate matches for human review, each pairing the submitted image with its published look-alike. The checks that operate at this level, from duplication within a manuscript to image and graph plagiarism against published sources, share one division of labor: software compares at scale, and people judge what it surfaces.
What a flagged match does and does not mean
A flagged match means two images are visually similar. It does not mean misconduct. Visual similarity has legitimate explanations that an editor rules in or out before anything else happens:
- The same authors reusing a methods or apparatus image across their own papers, with citation
- Licensed or permissioned republication, as in review articles that credit the original figure
- Earlier versions of the same work: preprints, conference proceedings, theses
- Shared instruments, calibration standards, or control materials that genuinely produce similar images
Screening experience points the same way. In her UKRIO expert interview, Jana Christopher describes the ratio of unintentional error to deliberate manipulation as difficult to know for sure, and a cross-paper match deserves the same presumption of an ordinary explanation until the authors have answered.
What a flagged match changes is the question an editor can ask. Without it, there is nothing to ask at all. With it, there is a specific, answerable query: this panel resembles this figure in this earlier paper; please explain the relationship and provide the original data. Presentation matters here: a screening platform that shows each match beside its candidate source, with the source identified, lets an editor resolve most flags in minutes and escalate the few that need attention. The decision stays human at every step. A database can say that two charts look alike; a person, with the authors' response in hand, decides what that similarity means. That judgment is easier with the base rate in mind, since most flagged images are honest preparation errors.
How Octym helps
Octym screens the figures, graphs, and charts in a submission against previously published material and surfaces suspected matches as signals, each shown beside its candidate source so the evidence is traceable. Powered by Proofig, the largest image-plagiarism database in the industry, the comparison treats graphs and charts as images from the start. Octym never issues a verdict: an editor reviews each match in context and decides what it means. See how this fits editorial screening in the solution for publishers.