Queensland University of Technology biostatistician Adrian Barnett just ran a machine learning filter across 2.6 million cancer studies and flagged 261,245 of them, roughly one in ten, as probable fakes. The BERT-based system trained on known retracted paper-mill work screened decades of research from 1999 to 2024 and found something sobering: the share of suspicious papers climbed from 1% in the early 2000s to over 16% by 2022.
What started as scattered academic fraud has become industrial. Paper mills now generate studies on assembly-line schedules, flooding journals with text that mimics real research but carries fabricated data. Peer reviewers miss most of it. Editors miss most of it. Human readers have no chance. So now three journals are testing whether an AI trained on previous fakes can spot new ones before they land in the literature.
The BMJ study frames this as an arms race. Barnett calls it a "scientific spam filter," which captures both the scale and the futility. Every time the filter catches a pattern, the fakers tweak their templates. Every template tweak forces the AI to retrain. Meanwhile, thousands of false papers slip through each year, citing each other, building fake citation chains, contaminating meta-analyses and systematic reviews that clinicians actually use to guide treatment decisions.
The problem hits hardest in cancer research because the field attracts money, draws attention, and involves high-stakes treatment protocols. A fabricated oncology study doesn't just sit in a database. It gets cited. It shapes clinical guidelines. It influences which drugs doctors prescribe.
Three journals are already piloting the screening tech. The real test comes when paper mills adapt faster than the AI learns, which history suggests they will. The fight has shifted from human editors versus fraud factories to one algorithm against another.
This article is informational and does not constitute medical, scientific, or investment advice.


