One of the roundtables at this year’s European Health Psychology Society conference was about constructs that may not exist. Not in a provocative sense, and not as a panel of contrarians. Just a group of researchers asking, out loud and in a scheduled session, how anyone knows that the thing they have spent a career measuring is a thing.

I have been carrying that question around since. It sits badly with the rest of the job.

The scale that survived

In 2020, Ian Hussey and Sean Hughes published an analysis with the honest title “Hidden Invalidity Among 15 Commonly Used Measures”. They took 26 scales, drawing on 144,496 test sessions from 81,986 participants, and ran each one through four validity tests rather than the one that usually gets reported.

On internal consistency alone, the measure most papers cite and most reviewers accept, 88 percent looked fine. When all four tests were applied together, 4 percent passed. Of the fifteen widely used measures in the headline analysis, one came through everything intact, and it was the Need for Cognition scale, which measures how much a person enjoys thinking hard about things. There is a joke in there that I am going to leave alone.

The result is contested, and the objection is worth stating because it is the interesting part. Whether a scale “passes” depends on where the authors set their cutoffs, and different reasonable cutoffs give different survival rates. Which is precisely the phenomenon the paper is describing. The field has no agreed threshold for when a measurement instrument is good enough, so the question of how many instruments are good enough cannot be answered without making the same discretionary call that created the problem.

Building a new one is publishable, checking an old one is not

What the field does instead is build more. A 2025 analysis in “Advances in Methods and Practices in Psychological Science”, titled “A Fragmented Field”, went through PsycTests, the main archive of psychological measures, and found 73,646 records. Of the measures introduced between 1993 and 2022, 56 percent show no traceable reuse by anyone, ever. That figure comes from automated name matching, so it means no reuse the method could find rather than proof of none, but the direction is not in doubt. Most measures are used once, by the people who made them, and then join the pile.

The same analysis found 1,230 measures that share a name with at least one other measure. There are nineteen separate instruments all called a theory of planned behaviour questionnaire. Two studies can therefore report results from “the” theory of planned behaviour questionnaire and have measured different things, and a reader has no way of knowing without chasing the citation back.

Constructing a scale is a publication. Validating someone else’s is a favour. The incentive points one way, and the archive is what that incentive looks like after thirty years.

Grit, and what happens when nobody checks

The clearest single case is grit, because it went from a construct to a bestseller to a school curriculum before the measurement question caught up with it. When Marcus Credé and colleagues ran a meta-analysis in 2017, they found grit correlating with conscientiousness at .84. To put that number in context, two established measures of conscientiousness correlate with each other at around .63. Grit was more strongly related to conscientiousness than conscientiousness was to itself.

That is not the same as saying grit is nothing. The perseverance facet does add a small amount of predictive value over conscientiousness alone, and the honest summary is that a well-marketed construct turned out to be mostly a rebranding of one that already existed, with a residue that might be real. But the reason it took a decade to find out is structural. Nobody gets a grant for checking.

The word depression is doing a lot of work

If this all sounds like an argument about obscure instruments, consider the one construct everybody believes in. Eiko Fried went through seven standard depression questionnaires and counted the symptoms they contain between them. There are 52. Twelve percent of those symptoms appear on all seven scales. Forty percent appear on only one. The average overlap between any two of the scales, measured as a Jaccard coefficient, is 0.41 — Fried issued a corrigendum in 2018 after a reader found a coding error in the original 0.36 figure, and the corrected number is the one that should now be cited.

Coding symptom descriptors into distinct categories involves judgement calls, so 52 is a considered count rather than a natural fact. The consequence still holds. Two people can both be classified as depressed while sharing almost no symptoms, and two studies of depression can be studying substantially different populations while using the same word in the title.

One response to this is Denny Borsboom’s network theory, which proposes that mental disorders are not underlying entities producing symptoms but causal systems made of the symptoms themselves: insomnia causing fatigue causing withdrawal causing low mood. I find it a more honest description of what the data look like. It is also routinely misread. Borsboom is arguing that disorders are real as systems, not that they are unreal, and the distance between those two readings is one careless sentence.

The aggression task with 157 scoring rules

My favourite example, if that is the word, comes from aggression research. The Competitive Reaction Time Task has participants compete against an opponent and, when they win, blast that opponent with a noise. The aggression score is built from the volume and the duration.

Malte Elson and colleagues went looking for how researchers actually calculate that score, and reported in 2014 that they had found at least thirteen distinct quantification strategies. The live database they have maintained since now lists 157 strategies across 130 publications, a figure that grows, so it needs date-stamping whenever it is quoted. More ways of scoring the task than papers using it.

Then they did the part that matters. They took three real datasets and re-analysed each one under different published scoring rules. The conclusions changed. Same participants, same button presses, same noise blasts, different answer about whether the effect was there.

What the paper count is actually measuring

Set this beside the volume. Indexed articles rose from roughly 1.92 million in 2016 to 2.82 million in 2022. Ioannidis and colleagues, writing in Nature in 2018, identified 265 authors who published more than 72 papers in a single year outside physics, which works out at one every five days.

Two figures usually get quoted alongside those and both are wrong.

Half of all published papers are not never cited: that is a 1990 artefact from counting letters and meeting abstracts, and the real ten-year uncited rate is around 4 percent in biomedicine. And the thousands of scientists publishing every five days is the physics-inclusive number, inflated by collaboration authorship conventions. I keep wanting to use them, which is a reasonable proxy for how much I want the argument to be true.

The argument does not need them. A field that has produced 73,646 ways to measure things, more than half of them used once, and that cannot say with confidence whether fifteen of its most common instruments measure what they claim, has an output problem that no citation statistic makes worse.

What the roundtable was really asking

The question on the table in Pafos was not whether psychology is a science. That framing is a waste of a good afternoon and no one in academia argues about that at this point.

The question was narrower and much harder to escape: what would it take to establish that a construct exists, and does anything in the current reward structure make anyone do it?

Fried, Flake and Robinaugh set out what the alternative looks like in a 2022 review, and none of it is mysterious. Specify the construct before measuring it. Report all four kinds of validity, not the flattering one. Reuse existing instruments instead of building a nineteenth version. Treat measurement work as a contribution rather than as overhead.

What struck me, sitting in that session, was how completely the profession already agrees with all of this and how little any of it changes what gets funded. I went home and looked at my own measures, which is the part of this I would rather not write down. Two of the scales I use routinely are on the list.