Should a paper that provides novel scientific insights and describes a significant scientific advance be rejected because it was entirely (or mostly) written by an AI agent? Should a mediocre paper that does not advance scientific knowledge in any significant way be accepted because it was written entirely by a human? The scientific community’s answer to both questions, judging by relatively recent trends in scientific publishing, appears to be yes.
Just a few days before I wrote this piece, one of the editors in chief of a machine learning journal made waves after posting about an experiment they did (https://medium.com/@TmlrOrg/asking-authors-about-their-own-papers-3d2e04e5dee0). Authors of several papers that would have been rejected from the journal were interviewed to find out if they really understood the contents of the paper they submitted. A large fraction of the authors could not. The editor’s first takeaway was that AI may have generated the papers and the authors had not verified the output before submission. In other words, what mattered most was whether a human could be held responsible for the content of the paper, not whether the content was valid. To be fair, the points the editor made were broader, addressing other side-effects of the widespread availability of generative AI (such as the substantial increase in workloads for editors and reviewers), and, at least for these papers, the quality was already assessed to be poor. However, the response of conferences and journals to the emergence of effective AI tools has been almost exclusively focused on the attribution question – is AI allowed to be an author without human supervision (with the consensus being a strong NO).
What about quality? How do we determine which papers are worth publishing? Since the emergence of open access publishing, the scientific community has increasingly veered towards the answer “everything”.
The journal PLoS One was created specifically as a counterpoint to the careful (and lengthy) scrutiny received by papers submitted for publication in “top” journals (I am using quotes because I personally do not believe a value ranking of journals is possible). The scientific output of humans deserves to be published (the fact that humans were involved was implicit at that time) whether or not the work represents a significant scientific advance (this article from Alice Robb summarizes well these claims) Even today, PLoS One advertises that they “evaluate research on scientific validity, strong methodology, and high ethical standards—not perceived significance.” (https://journals.plos.org/plosone/static/publish) This trend rapidly spread beyond PLoS both through new journals and publishers (e.g., PeerJ claims their evaluation is based on an “objective determination of scientific and methodological soundness, not on subjective determinations of ‘impact,’ ‘novelty’ or ‘interest’.”) and was even adopted by “top” publishers (e.g., Scientific Reports from the Nature Publishing group). The extreme outcome of this trend has been the emergence of preprint servers, such as Arxiv, which forego the review process altogether. Anything is publishable – period. What’s meaningful will be sorted out eventually by the readers.
This trend of moving away from expert quality control has largely become the “way to do business” in the scientific community. Notably, the Howard Hughes Medical Institute – one of the world’s most prominent biomedical research organizations – has recently stopped evaluating scientists based on whether their work has been published in peer-review journals. Importantly, this trend has not been embraced without resistance from some scientists, including myself. In the age of rampant misinformation, enabling the unchecked spread of unvalidated scientific claims is irresponsible.
And then, generative AI arrived. Suddenly, scientific papers that were (or appeared to be) scientifically sound but did not substantially advance scientific knowledge could be generated en masse with the push of a button. This led to the explosion in submissions to journals and conferences, explosion that led to the case study I mentioned at the beginning of this piece. The rapid AI-enabled generation of scientific papers challenged the conceptual foundation of the trend started by PLoS One. And how did the community respond? Not by rethinking the original premise that “anything is publishable”, but by adjusting it: “Anything is publishable by humans“. Journals and conferences have created AI policies focused not on the scientific quality of the work, but on whether human or machine was the author. And preprint servers, such as arXiv, have tightened their rules to prevent AI-generated “slop”. How about human-generated slop? That seems to still be acceptable.
To summarize, I believe we have lost perspective of the things that matter most, and generative AI is simply highlighting existing problems with the scientific publishing ecosystem rather than creating new ones. The welfare of humanity critically depends on the scientific progress. And scientific progress can be slowed and even reversed by large volumes of incorrect and/or uninformative results, whether generated by humans or machines. Peer review systems were developed over centuries specifically in order to remove the dirt from the cogs of the scientific machine, allowing it to more effectively and efficiently serve the needs of mankind. Adequate peer-review is a necessity, not an inconvenience, and I hope the challenges created by AI systems will prompt us to go back to the drawing board and create better and more resilient processes for conducting peer review despite attacks on this system from both scientists and politicians.