Epidemiologic Methods

Paper Mills Are Poisoning Science

Fake papers clutter journals, poison reviews and guidelines, reward bad incentives, and leave trash in the scientific record.

Read the essay ↓
A sign in a window of a store AI-generated content may be incorrect.

You can now pay a few hundred bucks to publish your name on a scientific study. No lab work, no peer review, no idea what the study’s about, no problem. Welcome to academia’s increasingly poorly kept secret, paper mills. Through these corrosive groups people are buying their slots on ‘indexed’ publications the same way you can buy fake followers on social media.

The scientific literature, our shared memory of humanity’s attempts to understand itself and the world around us, is being spammed with ghostwritten junk articles sold in bulk by shady brokers and rubber-stamped by complicit editors. If science is a public good, which I like to think it is, these guys are poisoning the well.

And to those thinking, “this is probably just some fringe grift on the outskirts of science”, think again. A new paper in PNAS titled “The entities enabling scientific fraud at scale are large, resilient, and growing rapidly” shows that this goes beyond a handful of rogue researchers fudging data a la Dan Ariely. The paper describes an ecosystem of coordinated fraud operating at a scale that could rival the real thing. The worst part? It’s going to be incredibly difficult to stop it.

The Paper

In what could turn out to be one of the most consequential meta-science papers of this decade, Richardson and colleagues set out to see how scientific fraud actually works as an industry. With massive datasets taken from PubMed, OpenAlex, Retraction Watch, and PubPeer, the researchers traced the footprints of thousands of bogus papers, zeroing in on the editors who greenlit them. They uncovered networks of collusion as dense as the Thule fog up at UC Davis on a cold winter day.

Here's what they found:

- A tiny fraction of editors are responsible for a disproportionate share of retractions. In PLOS ONE 0.25% of editors handled 30% of the retracted papers

- These editors often sent submissions to each other, like a corrupt little peer review cartel.

- Authors game the system by targeting these editors with submissions. Statistically speaking, these weren’t random assignments.

- They mapped networks of articles sharing duplicated images (like the same Western blot being used across multiple, unrelated studies).

- This suggests mass production with templated content.

- One outfit known as the Academic Research and Development Association (ARDA) lists the journals where it can “guarantee” publication

- When journals get deindexed, ARDA swaps in a new one. They call this “journal hopping.”

- Low-scrutiny, high-ouput areas like IncRNA and miRNA studies had retraction rates 20 to 40 times higher than higher-scrutiny fields like CRISPR research.

- Paper mill products are growing exponentially, doubling every 1.5 years. Real science doubles much slower, because it’s often a slow process.

Fig. 5. from Richardson et al. (caption from paper) Articles of fraudulent provenance have an apparent growth rategreater than that of the entire scientific enterprise and already far outpacethe scope of science integrity measures currently in use. (A) Retractions areincreasingly published in batches. The ∼2010 spike in the number of large-batch retractions is almost entirely attributable to a large swath of conferenceproceedings articles retracted by IEEE. For the first time since this spike, themajority of 2023 retractions were reported in batches larger than 10 articles.(B) Annual global scientific activity as measured by items labeled as “journalarticle” or “conference proceeding article” in OpenAlex (47), as retractedarticles reported by Retraction Watch, as PubPeer-commented articles and assuspected paper mill products. We make use of the linear trends observablein the log–linear plot to extrapolate these observations for the period 2020–2030. We show the 95% CI using shaded bands. The number of suspectedpaper mill products shows the largest growth rate, with a doubling time of1.5 y. (C) Annual global scientific activity captured by WoS as measured by thenumber of actively publishing journals, the number of journals deindexedannually by WoS, the number of journals with retractions, the number ofjournals with PubPeer comments, and the number of journals with suspectedpaper mill products. It is visually apparent that deindexing now occurs at alevel far below the level of occurrence of journals publishing suspected papermill products. These patterns hold for Scopus and MEDLINE (SI Appendix,Figs. S21–S23)

Subscribe now

But what is a paper mill?

A paper mill is exactly what it sounds like, a factory for churning out fake or low-quality scientific papers at scale. Think something like a college that was deemed a diploma mill, but instead of selling fake degrees, they sell authorship on peer-reviewed articles.

In part, they exist to exploit the pressure scientists face in a field where it’s publish or perish. In countries where licensing agencies require publication in an indexed journal, there’s a shadow industry lying in wait to exploit the system and offer a shortcut. “Just pay us and we’ll get your name on a paper!” (never mentioning that the paper might be plagiarized, fabricated, nonsensical, or stitched together from stock images and academic jargon). What matters for those buying is that it passes through a real journal and ends up with a DOI and line on the CV.

Perceived legitimacy is the product, and overworked and stressed (or straight up bad actors) are the customers. It’s real-seeming science, in real-seeming journals, indexed in legit databases like Scopus, PubMed, and Web of Science. And because the incentives can be skewed, favoring quantity over quality, speed over substance, there are always editors, journals, and brokers ready to make it happen.

Some of the mills ghostwrite entire papers and just slot people in as authors. Others act more like a middleman, brokering connections between desperate academics and complicit editors. A few, like ARDA, the one profiled most heavily in the PNAS paper, even go public facing with their websites advertising their services. When one of their preferred journals gets flagged or deindexed, it’s on to the next. It’s like the shady bar that gets shut down by the health department repeatedly, only to open under a new name. Same rats, new signage.

Lest you think this is just an obscure problem in low quality journals, these papers are showing up in decent journals with respectable impact factors and are being cited and make it into meta-analyses (giving garbage in garbage out a whole new look). It’s tempting to see them as somehow separate from science, but they’re clearly inside the system and using its own structures as the laundering mechanism.

So, let’s discuss some of the enablers.

Meet the enablers.

It turns out you don’t need a criminal mastermind for these kinds of things to happen. A couple of lazy editors, a “helpful” broker, and an under-resourced or indifferent journal will suffice. Most damning is the fact that this works because the system lets it.

Take PLOS ONE. It started as a bold experiment in open-access publishing but has become the publishing equivalent of a decaying, abandoned mall (sorry Puente Hills). Still technically functioning but now being used by people trying to do things they shouldn’t be. The paper found that a quarter of a percent of editors who edited 1.3% of articles were responsible for nearly a third of the retractions. More than half of whom had had articles of their own retracted in PLOS ONE (25 of 45 identified editors). That clearly goes beyond any excuse of an overworked reviewer making honest mistakes. These editors had a habit of handing manuscripts off to each other, passing them around like party favors. It’s a conspiracy so simple you could draw the diagram on a cocktail napkin at the bar.

PLOS ONE isn’t alone in this. Hindawi journals (owned by Wiley) showed a similar type of rot. A few editors are greasing the wheels, and a few hundred papers slip through with retractions lagging years behind the publication date. By then the damage is done. The scientific system doesn’t have the fastest reacting internal immune responses. It’s slow and inadequate, like a diabetic wound that never quite heals all the way.

But the most cartoonishly bold operator in this whole mess is ARDA, with an unscrupulous business model so naked it belongs on OnlyFans. They’ve listed 188 unique journals on their site as available venues. Some of which (17 to be exact) are suspected to be “hijacked journals” that were once legitimate but have been taken over completely by the paper mill. The dumbest thing about their racket? The fact that papers published through them often didn’t even relate to the journal in question. Examples given are a paper about roasting hazelnuts appearing in a journal on HIV/AIDS and an article on malware detection showing up in a special education journal.

This gives the impression that the editors were ither asleep at the wheel, complicit, or overwhelmed. Many journals are staffed by a skeleton crew of very few full time staff, relying heavily on volunteer or part-time editors (many of which are academics juggling a hundred other tasks). So even some blatant fraud slips through the cracks.

Not a victimless crime

It’s easy to shrug this off. “So some mediocre or fake papers get published, who cares?” But the harm is concrete and cumulative damage to fields of research. Every fake paper that gets indexed bloats the literature, muddies the search results, slows down systematic and narrative reviews, and forces real researchers to sift through garbage in search of gold. It’s a waste of our time.

Even worse is that some of the papers end up cited and included in meta-analyses, they inform grant applications and impact the AI models being trained on the literature. All this happening while the trust in science by Americans has been slowly diminishing. “Peer-reviewed” garbage that could have zero truth to them could be hoisted up by ‘influencers’ as evidence against real science, fooling those unable to tell the difference.

And for all of the talk about retractions and deindexing, it’s likely that the reality is most of the stuff doesn’t get pulled. Of the nearly 30,000 suspected paper mill articles on OpenAlex, only 28% had been retracted and only 10% ended up in deindexed journals. Most will just sit in the scientific literature forever, quietly polluting the water supply.

Why it’s not being stopped

Stopping this kind of fraud would require someone to actually be in charge of doing so. Instead we’re stuck with underpaid editors, overwhelmed journals, and volunteer sleuths like Elizabeth Bik trying to hold back a flood with a sieve.

Retractions are slow and rare, and like I said only about a quarter of papers from these mills/brokers ever end up retracted. Journal deindexing is also slow and doesn’t solve the problem fully. The journal may be deindexed, but the papers are still out there floating around. Institutional action is what we’d hope for but can’t really count on. The people responsible often have the conflict of interest of retractions making the journal look bad. A misconduct investigation is bad PR, so people just shrug and move on.

When a system is built to assume good faith, it’s easy for fraudsters to exploit it. And they end up doing so faster than the system can adapt. The authors of the PNAS paper make it pretty clear that if we don’t build a better infrastructure with detection, investigation, and sanctions handled by independent parties with actual resources, the enshittification will just keep getting worse. And every year we wait, the fake articles start to outpace the real science.

The Edge of Epidemiology is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

Originally published on The Edge of Epidemiology on Substack.