Skip to content
OctymPowered By Proofig AI
statistics

The GRIM test: when reported statistics are mathematically impossible

Reported means can be mathematically impossible given the sample size. How the GRIM test works, what an inconsistency signals, and why it is not proof.

2026-06-02 · Octym · 7 min read

Reported statistics can be mathematically impossible because averages of whole-number data can take certain values and no others. If 28 people answer on an integer scale, every possible mean is a multiple of 1/28, so a reported mean of 5.19 cannot have come from that sample as described. The GRIM test checks published means against exactly this constraint.

What does the GRIM test check?

The GRIM test, short for granularity-related inconsistency of means, checks whether a reported mean is arithmetically achievable given the reported sample size. Nick Brown and James Heathers introduced it in Social Psychological and Personality Science as a way to evaluate published psychology articles, and the logic transfers to any field that averages integer data: Likert items, symptom counts, error tallies, numbers of correct trials.

The reasoning fits in one sentence. A mean is a sum divided by a sample size, and when every raw value is a whole number the sum is a whole number too, so the mean must equal some integer divided by n. With 28 participants, achievable means move in steps of 1/28, about 0.036 wide. A mean reported to two decimal places has to land on one of those steps once rounding is accounted for. Consider what 28 integer responses can actually produce near a reported value of 5.19:

Sum of the 28 scoresExact meanRounds to
1445.14295.14
1455.17865.18
1465.21435.21

No whole-number sum yields 5.19 or 5.20. If a paper reports that the mean of 28 integer ratings was 5.19, the paper is wrong somewhere: in the mean, in the sample size, or in the description of the measure. The test does have boundaries. It is most informative for small samples reported to two decimals; once n reaches 100, the steps between achievable means become as fine as the rounding itself, and every two-decimal value becomes possible.

What does an inconsistent mean actually mean?

An inconsistent mean is a signal that something in the reporting chain went wrong, and most of the plausible explanations are mundane. Anyone running the check should walk through the innocent readings before entertaining any other kind:

  • A typo. A transposed digit in a table survives copyediting easily, because nobody recomputes descriptive statistics by hand.
  • Rounding behavior. Some software truncates where other software rounds. A true mean of 5.1786 printed as 5.17 by truncation will fail a check that assumes conventional rounding, through no fault of the authors.
  • The n beside the mean is not the n behind it. Missing responses, excluded participants, or a subgroup analysis can make the effective sample smaller than the column header suggests, with no intent to mislead.
  • The measure is not integer-granular after all. A scale built by averaging three items steps in increments of 1/(3n) rather than 1/n, so a naive check will flag means that are perfectly legitimate.

Only after these are ruled out does the space of explanations narrow toward data that never existed as described. Even then, the test cannot say which explanation applies: it establishes impossibility, not intent. The proportionate response to a flag is a question to the authors and, where policy allows, a request for the underlying data. Treating a flag as an accusation gets the logic backwards. The test identifies numbers that need an explanation, and most explanations are boring.

How often do impossible statistics appear in print?

Impossible means are common enough to be a systemic issue rather than a curiosity. When Brown and Heathers applied the GRIM test to a sample of published psychology articles, around half of the articles that could be tested appeared to contain at least one reported mean that was mathematically inconsistent with its reported sample size. These were peer-reviewed papers in respectable journals, and the inconsistencies sat in plain view, checkable by anyone with the paper and a calculator.

That base rate lands harder against the background of the reproducibility conversation. More than 70% of the 1,576 researchers who answered a Nature survey said they had tried and failed to reproduce another scientist's experiments. More than half told the same survey that they had failed to reproduce one of their own. Reporting errors are one strand of that tangle: a result whose descriptive statistics are wrong in print cannot be rebuilt from the paper, whatever actually happened in the lab.

The honest reading of the GRIM findings is not that half the literature is fabricated. It is that routine statistical reporting is far less reliable than the polish of a published table suggests, because until recently nobody was checking. Reviewers rarely recompute descriptive statistics; the working assumption has been that the numbers are what they say they are. GRIM showed how much that assumption misses, using arithmetic that was available to everyone all along.

Recomputing statistics at review time

Recomputation is the class of automated check that GRIM belongs to: rather than trusting a reported number, derive what it must satisfy and verify that it does. GRIM needs nothing beyond the reported mean, the sample size, and the scale, which is how Brown and Heathers could test published articles from the outside, without access to any raw data. Several members of the same family run entirely from a manuscript's text and tables:

CheckWhat it recomputesWhat an inconsistency can signal
GRIM-style mean checksWhether each mean is achievable given n and the scaleTypo, wrong n, non-integer measure, misreported data
Test-statistic consistencyWhether a test statistic, its degrees of freedom, and its p-value agreeCopy-paste errors, selective rounding, deeper problems
Percentage checksWhether percentages match their counts and denominatorsRounding drift, subgroup confusion
Totals and subgroupsWhether subgroup counts sum to the stated totalSilent exclusions, table assembly errors

These checks are deterministic and cheap, which makes them natural to automate alongside the broader set of manuscript integrity checks rather than reserving them for papers that already look suspicious. The practical question for editorial and integrity teams is placement. Run at submission or triage on a review platform, an inconsistency comes back to the authors as an ordinary revision query, easy to resolve while the data are still at hand. The same inconsistency found after publication becomes a correction, an expression of concern, or a public dispute. The arithmetic is identical; the cost of the moment is not.

That placement argument is now a capacity argument too. More than 10,000 research papers were retracted in 2023, a record annual figure, according to Nature. Screening capacity, not screening ambition, is what limits most editorial offices. Deterministic checks are the part of the workload that scales without adding reviewer hours, which is why major publishers have moved them into shared infrastructure such as the STM Integrity Hub.

When a recomputation does not resolve after an author query, it stops being an arithmetic question and becomes a process question. COPE guidance sets out the route editors follow for handling concerns raised about a published article and the corrections that may follow, and using that route from the first query keeps the record defensible if the matter goes further. Recomputation is also a useful illustration of where automation stops, a boundary covered in what automated checks cannot see.

How Octym helps

Octym's statistical validation and consistency checks recompute reported statistics during manuscript review and surface suspected inconsistencies as signals, each tied to the exact location in the manuscript where the number appears. The platform does not conclude why a mean fails a check; it presents the evidence so a person can ask the right question. For integrity and investigation teams, that turns a hunch into a documented, checkable starting point rather than a verdict.

See Octym on any manuscript.

Contact us, or log in.