Skip to content
OctymPowered By Proofig AI
image integrity

AI-generated figures in science: can they be spotted?

Generative tools produce figures experts cannot reliably spot by eye. What AI-generated image detection can honestly promise, and what publishers do next.

2026-06-30 · Octym · 6 min read

Generative tools can now produce scientific-looking figures that experts struggle to tell apart from genuine data images, which is why publishers are building automated detection instead of relying on the eye, as Nature reported. Detection is possible, but it returns probability, not proof, and a person still decides.

Why do AI-generated figures fool trained reviewers?

Scientific figures are unusually easy targets for generative models, and the reason is structural rather than technical. A western blot is a handful of grey bands on a noisy background. A micrograph is a field of roughly similar shapes. Neither contains the kind of semantic detail that makes a fake photograph of a street scene fall apart under inspection, because there are no hands with six fingers, no misspelled signage, no impossible reflections.

Reviewers also arrive with the wrong prior. Peer review evolved to evaluate whether an experiment was well designed and whether its conclusions follow from its data. It did not evolve to ask whether the data image is a photograph of anything at all. A reviewer looking at a blot is assessing what the bands mean, not interrogating whether the bands were ever on a membrane.

Nature reported that specialists are developing automated detection precisely because manual review does not reliably catch AI-generated figures. That is a significant admission from a field whose image screening has historically depended on expert human eyes, and it marks a real change in the threat model: the older problem was an author altering a real image, while the newer one is an image with no underlying experiment at all.

The practical consequence for editorial offices is that "looks plausible" has stopped being evidence of anything. A figure that survives visual inspection has passed a test that generative tools were specifically optimised to pass.

What can detection honestly promise?

Detection of AI-generated imagery works on statistical traces rather than on meaning. Generative models leave characteristic regularities in noise distribution, texture, and frequency content that differ from the artifacts produced by real microscopes, cameras, and imaging sensors. A detector learns those differences and reports how consistent an image is with synthetic generation.

That mechanism sets hard limits on what any honest system can claim, and they are worth stating plainly:

  • The output is probabilistic. A detector reports a likelihood, not a fact. There is no signature that appears in every generated image and never in a real one.
  • Compression and processing degrade the signal. Heavy JPEG compression, resizing, and figure assembly all remove the fine-grained traces detectors depend on, so a legitimately processed image can look more suspicious and a carefully laundered synthetic one can look cleaner.
  • The models keep moving. Detection trained on one generation of tools weakens against the next, so this is a maintained capability rather than a solved problem.
  • A high score is not a finding of misconduct. It is a reason to ask the author for the raw acquisition file and the instrument metadata.

None of this makes detection useless. It makes it a screening signal, in the same category as a plagiarism-similarity score: valuable for directing attention, never sufficient on its own to conclude anything about a person. The decisive step in almost every real case is not the detector output but the request that follows it, because a genuine experiment has raw files, instrument settings, and dated acquisition records behind it, and a generated figure does not.

What are publishers actually doing about it?

Publishers have responded collectively rather than individually, which tells you something about how the problem is understood. Major publishers collaborate through the STM Integrity Hub on shared infrastructure for screening submissions, including paper-mill signals and other research-integrity concerns.

Shared infrastructure matters for synthetic figures specifically because the economics favour volume. A paper mill producing fabricated manuscripts at scale can generate figures far faster than any single journal can examine them, and the same fabricated image may be submitted to several journals in sequence. A publisher checking only its own submission stream sees one instance; a shared system sees the pattern.

At the individual desk, the response has been to move image screening earlier and make it routine. Between 20 and 35% of manuscripts checked at acceptance already raise an image-related query during editorial screening, according to image integrity analyst Jana Christopher speaking to UKRIO, while acceptance is ultimately rescinded for only 1 to 8% of the manuscripts she screened, so the large majority of flagged papers keep their acceptance. Synthetic figures arrive into a workflow that already exists, which is a genuine advantage: journals do not need to invent a process, only to extend one.

The context sharpening all of this is volume. More than 10,000 research papers were retracted in 2023, a record annual figure, according to Nature. Screening capacity, not screening ambition, is now the binding constraint at most editorial offices.

What should an editor do with a flagged figure?

An editor holding a suspected AI-generated figure should treat it as the opening of a documented process, not as a conclusion. COPE guidance gives editors a documented route for handling concerns raised about a published article and the corrections that may follow, and taking that route is what separates a defensible decision from an improvised one.

In practice the sequence is short. Ask the author for the original acquisition files, the instrument or software that produced them, and the acquisition dates. A real experiment answers that request quickly, because the files exist in a folder somewhere. A fabricated figure produces delay, substitution, or an explanation that does not survive a follow-up question.

Record what was checked and what was returned. This is the part editorial offices most often skip and most often regret, because a question raised about the paper two years later needs the contemporaneous record, and an integrity team reopening a case with nothing but a memory of a concern starts from a blank page. Teams handling this work need evidence they can hand to a committee, and that record has to be created at the moment of the query.

Keep the register neutral throughout. The author of a flagged figure is, statistically, far more likely to have made an error than to have fabricated data, and the correspondence should read as though that is true until the evidence says otherwise. That default is well founded: across image flags generally, most flagged papers keep their acceptance.

How Octym helps

Octym reviews figures for signs of AI generation alongside manipulation, duplication, and reuse, and returns each finding as a signal with its evidence and location attached. Findings are ranked so that limited editorial attention lands where it matters most. Nothing is presented as a verdict about an author, and no finding closes a question on its own: a person always decides what a flagged figure means and what to ask for next. The full list of checks is on the what we check page.

See Octym on any manuscript.

Contact us, or log in.