Skip to content
OctymPowered By Proofig AI
references

Hallucinated references: how fake citations get into real papers

Language models invent citations that look real, and they survive review. How hallucinated references reach published papers, and how to catch yours first.

2026-07-21 · Octym · 6 min read

Fake references get into real papers when authors let a language model draft or format citations and trust the output. Models generate plausible references to works that do not exist, co-authors and reviewers rarely resolve every entry, and fabricated citations pass into the published record unnoticed. Resolving each reference against an open scholarly index before submission catches them.

Why do language models fabricate citations?

Language models fabricate citations because they generate text rather than retrieve records. When an author asks a chat model for sources, the model produces strings that are statistically shaped like citations: author surnames active in the field, a title assembled from the vocabulary of the topic, a genuine journal name, a volume, a page range, a DOI in valid format. Nothing in that generation step consults a bibliographic database, so nothing anchors the output to a work that actually exists.

This behavior has been quantified in the peer-reviewed literature. A study in Scientific Reports found that a majority of the bibliographic citations generated by GPT-3.5 were fabricated. The same study found that GPT-4 did better yet still produced a substantial share of fabricated or erroneous citations (Scientific Reports). Model progress narrows the problem without removing it, because the underlying mechanism is unchanged: unless a tool performs an explicit lookup against a real index, its citations are predictions, not references.

The failure also travels beyond the obvious case of asking a chatbot for a reading list. Drafting assistants that suggest related work, summarizers that compress a literature review, and formatting helpers that "complete" a partial citation all pass text through the same generative machinery. A fabricated reference emerges looking exactly as confident as a real one, which is precisely why authors trust it.

How do fabricated references survive peer review?

Fabricated references survive peer review because no stage of the traditional pipeline resolves every entry in a bibliography. Reviewers are recruited to judge the science: methods, analysis, novelty, the strength of the claims. Most will check the handful of references they know well, perhaps the ones citing their own work, and pass over the rest. Editorial production then checks reference formatting against house style, a test that a machine-generated citation, built to look typical, passes comfortably.

Two properties make these entries hard to spot by eye. They are assembled from real ingredients, genuine journal names, plausible author combinations, well-formed DOI strings, so there is no visual tell. And fabrication shades into ordinary error. The Scientific Reports study found that even citations pointing to real works frequently contained substantive errors: wrong years, wrong volumes, wrong page numbers. A reviewer who notices a discrepancy has every reason to read it as a typo rather than as a symptom that the bibliography was generated and never checked.

Timing compounds the gap. Reference lists are often finalized late, under deadline, after the scientific argument has already been reviewed in draft. And once a fabricated citation is published, it can propagate: later authors copy references secondhand from papers they trust, so a single unresolved entry can echo through a citation chain for years.

What does resolving a reference actually mean?

Resolving a reference means taking the claim every citation makes, that a specific work by specific authors exists in a specific venue, and finding the record that proves it. Resolution is a mechanical lookup, and open infrastructure now makes it possible at scale. OpenAlex is a fully open index of hundreds of millions of scholarly works with their metadata and identifiers, which turns the existence question into something answerable by query rather than by trust. Retraction status is public as well: Crossref acquired the Retraction Watch database in September 2023 and made it openly and freely available as what Crossref calls the largest single open-source database of retractions.

A full resolution pass asks four questions in sequence:

Resolution stepQuestion it answersA failure suggests
DOI lookupDoes the identifier point to the claimed work?A fabricated or mistyped DOI
Index matchDoes an indexed record match the title, authors, and year?A hallucinated reference
Metadata comparisonDo the journal, volume, and pages agree with the record?Citation errors that break traceability
Retraction checkIs the work still part of the trusted record?A retracted reference carried forward

Each row produces a signal for a person to examine, since indexes have gaps and honest typos exist. Reference resolution sits alongside image, statistical, and plagiarism review in what a full manuscript check covers, and it is the specific layer that catches hallucinated citations.

How should you check your own bibliography before submission?

Checking your own bibliography before submission is a short, mechanical routine, and it is worth running on any manuscript where a model touched the references at any point.

  1. Resolve every DOI, not a sample. Fabricated entries do not cluster where you expect them, and a spot check misses the point of a mechanical error.
  2. Search each title in an open scholarly index. If an entry that claims to be a published article returns nothing in an index such as OpenAlex, treat it as unverified until you locate the actual work.
  3. Compare the full metadata, not just the title. Author order, year, venue, volume: the Scientific Reports finding that real works were frequently cited with substantive errors (Scientific Reports) means a matching title is not enough.
  4. Treat AI-touched entries as unverified by default. Whether the model drafted the citation, reformatted it, or completed a partial one, the entry needs a resolution pass before it is trustworthy.
  5. Check retraction status while you are there. A reference can be real and still be withdrawn, and the retraction record is openly available through Crossref.

Done by hand, this costs an evening for a typical paper. The cost of skipping it lands later and higher: a bibliography that fails resolution invites questions about everything else in the manuscript. Many researchers now fold reference resolution into a broader pre-submission integrity check so the whole bibliography is verified in one pass rather than entry by entry. A fabricated reference is one of two ways a bibliography fails; the other is a citation that was real when you wrote it and has since been retracted.

How Octym helps

Octym resolves the references in a manuscript against open scholarly records, backed by OpenAlex, and flags entries that fail to resolve or whose metadata does not match the indexed work. Every suspected finding is surfaced as a signal with its evidence traced to the source record, never as a verdict. The author or editor reviews each flag and decides what to fix. Reference resolution is one layer of what Octym checks across a manuscript before it reaches a reviewer.

See Octym on any manuscript.

Contact us, or log in.