You can use ChatGPT to improve your manuscript's language, structure, and clarity, and it does that well. You cannot use it as an integrity review: chatbots fabricate citations, answer confidently without evidence, cannot examine images, and pasting an unpublished manuscript into a consumer service raises confidentiality problems. Integrity review needs grounded, validated checks, with a person deciding.
What ChatGPT does well in manuscript preparation
A general-purpose chatbot is genuinely useful for manuscript preparation, and pretending otherwise would be dishonest. Used carefully, ChatGPT can tighten prose, restructure a rambling introduction, suggest clearer topic sentences, and flag places where an argument skips a step. For researchers writing in a second language, that assistance can level a playing field that has long favored native English speakers.
A chatbot also works as a rehearsal audience. Ask it to summarize your abstract in two sentences and you learn whether the abstract communicates what you think it does. Ask it to argue against your discussion section and it can expose weak reasoning before a reviewer does. These are drafting aids: the researcher stays the author and checks every suggestion.
Two boundaries apply even at this stage. One is disclosure: the ICMJE Recommendations require authors to disclose the use of AI-assisted technologies in manuscript preparation and state that chatbots cannot be listed as authors. Editing help is legitimate; undisclosed editing help is a problem. The other is accountability: whatever the chatbot writes, the authors remain responsible for it. A fluent paragraph containing a subtly wrong claim is worse than an awkward paragraph containing a correct one, because reviewers forgive awkwardness and do not forgive error.
None of this is manuscript review. It is writing support, and the distinction matters for everything that follows.
Why can't ChatGPT verify citations?
Citation checking is where the chatbot-as-reviewer idea collapses, because language models generate references the same way they generate everything else: by producing plausible text. A peer-reviewed study in Scientific Reports found that a majority of the bibliographic citations GPT-3.5 generated were fabricated, and that GPT-4 still produced a substantial share of fabricated or erroneous citations. The same Scientific Reports analysis found that even citations pointing to real works frequently contained substantive errors.
The failure mode matters more than the failure rate. A fabricated reference does not arrive flagged as uncertain; it arrives formatted like every real one, with a plausible author list, a real-sounding journal, and an identifier that resolves nowhere. And a chatbot asked to review your bibliography has no mechanism for looking anything up: it does not query a bibliographic database, does not resolve DOIs, and does not know whether a paper was retracted last month. When you ask whether your references are correct, it answers from the same statistical process that invents references to begin with.
Hallucination extends past the bibliography. Ask a chatbot whether your statistics are sound and it produces a confident, articulate assessment untethered to any recalculation. Ask whether your figures contain duplication and it answers without inspecting a pixel. There is no evidence trail behind any of it: no record you can open, no source you can check. The output is a review-shaped answer, not a review.
Can ChatGPT check images and figures?
Image integrity sits beyond what a chatbot can review, and generative AI is simultaneously making the problem harder. Nature news has reported that generative models can produce scientific-looking figures that experts struggle to distinguish from genuine data images. The same Nature news reporting notes that publishers and integrity specialists are developing automated detection tools precisely because manual review does not reliably catch these images.
Screening figures is forensic work: comparing regions within and across figures, matching images against previously published work, and examining artifacts that betray manipulation or generative origin. A text model responding in a chat window does none of this. It can describe what image manipulation is; it cannot examine your images. The gap between those two abilities is easy to miss, because the chatbot's description sounds expert.
Editors face the same asymmetry at scale, and the publishing industry's response has been specialized shared infrastructure rather than general chat tools. Through the STM Integrity Hub, major publishers collaborate on screening submissions for paper-mill signals and other research-integrity concerns. The direction of travel across the industry is purpose-built, validated detection with defined outputs, applied consistently across submissions. A conversational model that produces different answers to the same question on different days cannot anchor that kind of process.
The confidentiality problem with pasting manuscripts into a chatbot
Confidentiality is the risk researchers most often overlook when pasting a manuscript into a chatbot. An unpublished manuscript is confidential material: unpublished data, unprotected ideas, sometimes information restricted by collaborators, funders, or patient-privacy rules. A consumer chatbot is a third-party service, and what happens to pasted text is governed by that service's terms, which vary by provider, plan, and settings. Depending on configuration, submitted content may be retained or used to improve models. For work that is not yet published, that is a real loss of control.
The problem is sharper for editors and peer reviewers. A manuscript under review is entrusted to you; it is not yours to share. Uploading someone else's unpublished work to an external AI service risks breaching reviewer confidentiality, and many publishers now explicitly instruct reviewers not to do it. The same logic applies to editorial staff screening submissions: the manuscript's authors never agreed to have their work processed by a consumer service.
The practical question for any AI review tool is therefore also infrastructural: where is the manuscript processed, who can access it, how long is it retained, and is it ever used for training? Before adopting any service, read how it handles data security and confidentiality. Analysis on private, access-controlled infrastructure with clear retention terms is a different proposition from a free chat window, even when the two look similar on the surface.
What grounded manuscript review means
Grounded review means every finding is tied to checkable evidence, which is exactly what a chatbot cannot offer. A grounded citation flag shows the bibliographic record a reference did or did not resolve to. A grounded image finding shows the matched regions side by side. A grounded statistical flag shows the reported values and the inconsistency between them. The reviewer can open the evidence, disagree with it, and overrule it.
| Dimension | Consumer chatbot | Grounded integrity review |
|---|---|---|
| Language and structure feedback | Strong; its core competence | Not the goal |
| Citation checking | Generates plausible references; documented fabrication risk | Resolves references against bibliographic records |
| Image screening | Cannot inspect images | Validated detection across figures and prior publications |
| Evidence trail | None; fluent assertion | Each finding links to its source |
| Confidentiality | Consumer terms vary by provider and plan | Private, access-controlled processing |
| Final judgment | Sounds equally confident right or wrong | A person weighs the evidence and decides |
Validated matters as much as grounded: detection methods tested against known cases behave in characterized ways, while a prompt improvises fresh behavior every time. Grounded review also has defined scope, a documented set of checks spanning images, references, statistics, and methods, rather than an open-ended conversation that will attempt anything and verify nothing. That is the design premise of a dedicated review platform: each check built for one job, producing evidence a person can audit, with the human decision kept at the center. The clearest case study is the bibliography, where fabricated citations reach published papers precisely because they look correct.
How Octym helps
Octym takes the grounded path. Its checks run on private infrastructure and surface suspected issues as signals, each traced to its source: the matched image region, the resolved reference record, the specific inconsistency in the text. It never issues a verdict; a person reviews the evidence and decides what it means. The reasoning behind that design is laid out in why Octym.