Experiments fail to reproduce mostly because the published account is incomplete, not because the underlying science is wrong. Methods omit steps, reagents and analysis choices go unstated, samples are too small to give a stable answer, and analysis decisions get made after the data arrive. Another lab then rebuilds the protocol from guesswork and gets a different result.
How widespread is the reproducibility problem?
The reproducibility problem is common enough that most working scientists have hit it themselves. More than 70% of 1,576 researchers surveyed said they had tried and failed to reproduce another scientist's experiments, and more than half said they had failed to reproduce one of their own, in a 2016 survey published by Nature. That second finding is the one that reframes the argument: it describes researchers unable to repeat work they designed, ran, and know intimately.
That detail rules out the easy explanations. If a large share of a field cannot repeat its own experiments, the cause is not incompetence at the receiving end, and it is rarely misconduct. Something is being lost between the bench and the page. The account that reaches a reader is a compressed version of what actually happened: a methods section written last, trimmed to a word limit, describing a protocol the author has performed so many times that the parts they do automatically no longer feel worth writing down.
For a research office the aggregate matters more than any single case. Work that will not replicate absorbs grant funding, stalls follow-on projects, and consumes the time of every lab that tries to build on it. It also surfaces late, usually after publication and after another group has spent a year failing to reproduce the result. By then the options are limited and public. Institutions that treat reproducibility as a property of the manuscript, rather than a property of the field, get to intervene while the work is still theirs to fix.
Where does irreproducibility actually come from?
Irreproducibility comes mostly from what a paper leaves out and from choices made after the data arrived. Four causes account for a large share of failed replication attempts, and none of them require anyone to have behaved badly.
Under-specified methods head the list. A protocol described in summary form reads fine to a reviewer and is unrepeatable at the bench: incubation times, temperatures, passage numbers, and the order of steps go unstated because they are obvious to the person who ran them. Missing reagent detail compounds it. An antibody, cell line, plasmid, or software package named without a catalogue number, identifier, or version leaves the next lab guessing which of several similar products was used, and those products do not behave identically.
The other two causes sit in the analysis. Flexible analysis, where the choice of test, exclusion rule, or subgroup is settled after seeing the data, produces results that look conclusive and do not survive a second dataset. Underpowered designs produce estimates too unstable to replicate: a small study that reaches significance often does so by overstating the effect, so an honest replication lands closer to the truth and reads as a failure. Authors and PIs rarely experience any of this as a reporting problem, which is why it persists through drafting and into submission.
| Source of failure | How it shows up in the manuscript | What closes the gap |
|---|---|---|
| Under-specified methods | Steps summarized rather than specified; conditions a newcomer would need are missing | Protocol-level detail, or a deposited protocol the paper points to |
| Missing resource detail | Reagents, cell lines, strains, and software named without identifiers or versions | An identifier for every key resource used |
| Flexible analysis | Analytical choices presented as if fixed in advance, with no note of what was planned | An explicit split between preplanned and exploratory analyses |
| Underpowered design | Sample sizes reported without justification; effects sitting near the threshold | A stated sample-size rationale, set before data collection |
| Selective reporting | Only the conditions and outcomes that worked appear in the results | All measured conditions and outcomes reported, including null ones |
Why reporting guidelines are the practical lever
Reporting guidelines turn the question "is this study rigorous?" into "is this account complete?", and the second question can be answered before submission. The EQUATOR Network maintains a searchable library of reporting guidelines covering the main study types: CONSORT for randomized trials, ARRIVE for animal research, STROBE for observational studies, PRISMA for systematic reviews, and many more. Each one is a checklist of items an account of that study type has to contain, from how participants were allocated to how missing data were handled.
Completeness is also checkable arithmetically in places, which is easy to forget. Reported means can be tested against their sample sizes, and Brown and Heathers found that around half of the testable psychology articles they examined contained at least one mathematically inconsistent mean. A result that cannot be recomputed from the paper's own tables is one a reader has no route to reproduce.
The distinction matters because design and completeness fail at different moments. Once data collection ends, the design is fixed: an underpowered experiment cannot be made powerful during revision. Completeness behaves differently. Every item on a reporting checklist can still be added, clarified, or justified while the manuscript is a draft, and the EQUATOR library exists so that authors do not have to guess which items their field expects to see.
Funders have pushed in the same direction. NIH applications are expected to address rigor and reproducibility explicitly, and European programs including ERC and Horizon Europe attach their own conditions on reporting and data availability. The practical effect is that a methods section now serves three readers at once: the reviewer judging the claim, the funder checking that stated commitments were honored, and the scientist who will try to repeat the work. Completeness is what all three need, and it is the dimension of rigor a check before submission can still change.
What can institutions do at the manuscript stage?
Institutions can act at the point where a manuscript leaves the building, the last moment the work is still theirs to correct. The steps are unglamorous and mostly procedural:
- Match the manuscript to its guideline. Identify which reporting checklist applies to the study design, then read the methods against it before submission rather than after a reviewer does.
- Require identifiers for key resources. Catalogue numbers, resource identifiers, strain designations, and software versions cost minutes to add and are close to impossible to reconstruct years later.
- Ask for the sample-size rationale in writing. If nobody can state how the number was chosen, that is worth knowing before the paper is public.
- Separate preplanned from exploratory analysis. Labeling exploratory results as exploratory costs nothing and protects the claim that matters.
- Check internal consistency. Numbers in the text that disagree with the figures, or legends that describe panels the figure does not show, travel with the same manuscript as everything else a manuscript check examines.
There is a downstream reason to move earlier as well. The screening layer at journals keeps hardening: through the STM Integrity Hub, major publishers collaborate on shared infrastructure that screens submissions for paper-mill signals and other research-integrity concerns. Issues that once surfaced in peer review, if they surfaced at all, are now raised by automated checks at the door, and what those checks find goes to an editor rather than to the author. An institution running its own review before submission sees the same class of issue while it can still be handled quietly, in a draft, by the people who did the work. The same expectations now start earlier, at the application stage, where funders ask applicants to address rigor explicitly.
How Octym helps
Octym reviews methods and reproducibility reporting before a manuscript is submitted, as part of a single pass over the paper. Where a described method omits detail another lab would need, or an analysis leaves its choices unstated, the gap is surfaced as a signal with the supporting evidence, located in the manuscript and traced to its source. Octym issues no verdicts: the author, the PI, or the research office decides what to add, clarify, or leave as written. That gives research offices a clear look at the work while it is still theirs to change.