For the past two years, part of my work as a psychology PhD researcher has meant reading other people’s data before I read my own. Not out of suspicion, most of the time. Just the ordinary, unglamorous discipline of checking whether a sample size holds up, whether an effect looks too clean, whether the numbers in a results table actually add up to the story the abstract is telling. Most of the time, they do.
Lately, I have started reading a little more carefully than I used to.
In the early 2000s, something close to one in a hundred cancer research papers already carried the textual fingerprint of a paper mill: an operation that manufactures scientific studies at scale and sells authorship slots to researchers under pressure to publish, with little regard for whether the underlying data exists at all. By 2022, according to a BMJ study published earlier this year, that share had risen to just over 15 percent of the field’s annual output, 26,457 of 171,656 papers screened in that single year.
Not quite one in six. Close enough that the distance stopped feeling like a rounding error and started feeling like a trend line with somewhere further to go.
What a paper mill actually sells
The term sounds almost quaint, like something out of an industrial-era novel. In practice, a paper mill is a business. It sells fabricated or lightly disguised research to academics who need a line on their CV and cannot, for whatever reason, produce the work themselves: overloaded early-career researchers, people in systems where a fixed number of publications is a condition of graduating or being promoted, sometimes people who simply found a shortcut and took it. Adrian Barnett, the biostatistician at Queensland University of Technology who led the BMJ study alongside lead author Baptiste Scancar and co-author Professor Jennifer Byrne of the University of Sydney, put it plainly: paper mills “are producing ‘research’ on an industrial scale, and our findings suggest the problem in cancer research is far larger than most people realised.”
Byrne has spent years cataloguing the specific tells: gene names substituted for one another like mismatched puzzle pieces, experiments that describe procedures no lab actually ran, results sections that could be swapped between unrelated papers without anyone noticing. What Barnett’s team added was scale. Instead of one editor squinting at one suspicious paper, they trained a language model on 2,202 confirmed paper mill papers already pulled from the literature, then ran it across 2.6 million cancer studies published between 1999 and 2024.
Building a spam filter for a body of research
Barnett described the tool bluntly: “We’ve essentially built a scientific spam filter. Just like your email system can spot unwanted messages, our tool flags papers that match the writing style and structure we see in retracted, fraudulent work.” The model caught the pattern because paper mills, for all their industrial ambition, are still running a production line: “Most likely, they’re relying on boilerplate templates which can be detected by large language models that analyse patterns in texts,” he noted.
There is something almost funny about that, in the way that the truest things about fraud usually are. The mills exist because fabricating research at scale requires templates, and templates are precisely the repetitive structure a model learns to recognize. The same industrial logic that let the problem grow past one editor’s capacity to catch it is what finally let a machine catch up to it.
The same tool on both sides
There is an obvious irony sitting inside all of this that Barnett’s team doesn’t shy away from: the technology accelerating the fraud and the technology catching it are increasingly the same technology. Large language models make it faster than ever to generate a plausible-sounding methods section, a results paragraph shaped like a hundred others, an abstract that reads like real research without describing anything that actually happened. And a large language model, trained the right way, is also what finally made it possible to flag a quarter of a million suspect papers in one pass instead of one editor’s lifetime of manual review. Neither side of that equation is going away. If anything, the arms race probably just started.
The part that is harder to automate
I don’t work in cancer research. My own field runs on a smaller, stranger version of the same incentive structure: a fixed number of publications expected before you can call a degree finished, journals that reward a clean story over a complicated one, a quiet pressure to produce something before you fully understand what you have found. I have sat in seminar rooms where the unspoken question was never “is this true” but “is this enough,” and I don’t think that distinction is unique to psychology, or to me.
That is the part a spam filter cannot fix. You can catch a fabricated paper. You cannot as easily catch the researcher who cut a real corner because the alternative was not finishing at all, or the reviewer who let something through because they were reviewing five other manuscripts that week for free. The paper mill is the visible symptom. The system that makes manufactured research a rational shortcut for enough people to sustain an industry is the part that stays largely untouched by any single tool, however good.
Why the number matters more than it sounds
“Cancer research influences clinical trials, drug development and patient care. If fabricated studies make their way into the evidence base, they can mislead real scientists and ultimately slow progress for patients.”
That is Barnett again, and it is the sentence that stayed with me longest, because it names the actual cost. A fabricated paper in a low-stakes corner of the literature is an embarrassment. A fabricated paper on a cancer mechanism that gets cited by researchers building the next study, or referenced in a systematic review, or folded into the background section of a grant proposal, becomes something closer to sediment: it settles into the foundation of the work that comes after it, and nobody downstream necessarily knows it is there.
I don’t have a tidy fix to offer, and I am wary of anyone in my position who claims to. What I keep coming back to instead is smaller and less satisfying: the number climbing from one in a hundred to nearly one in six is not really a story about a handful of bad actors. It is a story about how many of us are working inside systems that made cutting corners look, for a while, like the more sensible option. I read my own field’s papers a little differently now. Not because I trust anyone less, exactly.
Just because I’ve started to notice how much of trust in science was always, quietly, doing more work than it should have had to.