The problems
The way how we do scientific research is changing rapidly. The number of scientific manuscripts submitted to the journals and also to the preprint server arxiv.org has been growing rapidly, and this was increased by the appearance of capaple AI models.
Human referees can not keep up with this increased workload.
Various suggestions are starting to appear in the publishing companies. At the same time, referees started to experiment with AI tools, whether allowed or not. There is a clear need to regulate AI use in both producing and reviewing scientific work, and such guidelines are being formulated in these months.
As an example, APS recently updated its guidelines for AI use, for authors and referees. The new guidelines are more liberal than before, they do allow the use of AI, although with some limitations.
One such limitation that is present practically everywhere is that the manuscript should not be uploaded to an unrestricted AI tool. The reason for this strict limitation is clear: the manuscript is in principle confidential, and the authors may want to avoid using their material for training of new LLM’s, before the publication process is finalized.
This is the point where I strongly disagree with the current guidelines. I have a different proposal. I think it could be realistically applied in theoretical physics, and maybe also in some other fields.
The proposed alternative guideline
The key is the arxiv preprint server: If the authors acknowledge that they had uploaded their manuscript (or an earlier version of that) to arxiv prior to sending the material to a journal, then referees should be allowed to upload the arxiv version to any unrestricted AI tool.
Uploading to arxiv means that the document became openly searchable and readable, by humans and machines alike. Since the time of uploading it is possible that the material was used to train new LLM’s. There is typically a delay between the appearance on the arxiv, sending the material to a journal, selecting the referees, and the actual start of the refereeing process. It happens often that this delay spans a few months. By that time it is likely that at least one AI company used that particular document as training data.
It is possible that some authors do not wish to upload their material to arxiv. This happens sometimes with distinguished journals, special type of research papers. It might happen nowadays more often. In such cases I support the original guidelines. A lack of uploading to the arxiv signals the intention of the authors to not announce their results before the publishing process has been finalized. This should be respected.
If there is a difference between the arxiv version and the version sent to the journal, this could be communicated to the editors and to the referee, who should take this into account while preparing the report.
Potential advantages
In my experience present day frontier models do an excellent job in reviewing research articles. They can check all the analytic computations presented in the paper, and they can also reproduce some part of the numerical examples, at least those of some warm up examples. They can spot minor issues such as inconsistencies in notations, mismatches in figure captions, etc, but they also spot major issues. They are able to figure out if the scientific argument is not sound, if there are other potential explanations. They also spot omissions in the list of references.
This experience that I describe comes from running adversarial reviews on my own papers and on arxiv versions of some of the work of other researchers. My general impression is that the best models today are a better, more thorough referee than myself. Maybe I am not the best referee, but also I don’t think I am the worst. Therefore, it is likely that AI as a standalone agent is better than the median human referee. Also, AI will perform many checks that physicist referees simply just don’t do: complicated checks of the computations, and even of the numerics occasionally.
Furthermore, there is a clear gain in time. An AI assisted process is much faster, and the gain is likely very similar to what we gain on the side of producing the articles. So if AI use allows the authors to write x times more articles per month, then AI use allows the referees to review x times more articles per month. This could be one way to deal with the increased workload.
Therefore, I strongly suggest to use AI assistance in the refereeing process.
At the same time, the process should not be completely autonomous. At least, not yet, not with the models that we have today.
In my experience the human referee will have observations that the AI missed. Therefore, a collaboration between the AI and the human can work like using multiple experts for a single review.
Currently the most problematic part of the review process is judging the novelty of a research article, and whether or not it should appear in a particular journal. I understand that humans prefer that AI should not have a say in the decision. Judging such a complex question may indeed be more difficult for an AI, certainly more difficult than reproducing a series of analytic computations in the paper. After all, this involves a lot of human elements. There might come a time when we can trust an AI even with this part of the job, but perhaps not today.
Therefore, it is absolutely necessary that the human referee stays in the loop. At least for now.
The future far ahead
No one really knows how science will look like in 5 or 10 years. Will the singularity happen? What will it bring us?
I prefer not to speculate about this. Other people are doing it anyways. However, I have a short note.
It might be possible, that the arena of announcing scientific results will gradually shift from preprints and journal articles to other platforms, such as blogs, X posts, etc. We are seeing this happening.
It might be desirable to keep the existing structures, including preprint servers and peer review. But we can not close our eyes on the shifts that are actually happening today.

