The Association for the Advancement of Artificial Intelligence (AAAI) received nearly 29,000 paper submissions to its 2026 conference—approximately double the number for the previous year’s edition. These papers were from 75,000 unique authors, and the size of the program committee overseeing peer review tripled to 28,000 members compared to the previous year. A similar trend has been seen in many academic conferences and journals.
Scholarly manuscripts seeking to advance the frontier of human knowledge must be peer-reviewed by independent experts before their ideas can enter scholarly discourse. In other words, peer review is the cornerstone of academic research. It is what lends trustworthiness to published science.
Since the dawn of digital and then digital-only academic journals in the 1990s, the volume of published peer-reviewed scholarly articles has ballooned. So has the number of articles that get submitted for review at journals and conferences. The latest deluge of submissions, as witnessed by the organizers of the AAAI conference, is coincident with the rapid advancements in large language models (LLMs) with increasingly impressive reasoning and genre-appropriate writing skills.
Much peer review is conducted by early-career researchers—primarily graduate students, postdoctoral scholars, and junior professors. In nearly all instances, it is conducted by scholars without monetary compensation, as a voluntary service to their discipline and profession.
Given that early-career scholars are already overburdened, the temptation to automate peer review duties is real. Calls for automating aspects of peer review have grown louder in the last few years. Publishers, editors, and conference chairs have increasingly argued that the use of AI would speed up the process, relieve the burden on reviewers, and make peer review more reliable, consistent, unbiased, and accurate.
More than half of researchers recently surveyed reported using AI tools while reviewing papers. But efforts to automate various aspects of peer review, such as checking for plagiarism and methodological flaws, are nothing new—they’ve been around at least since the 2010s.
The question is whether AI can be trusted for peer review.
As an information scientist I study trustworthiness of generative AI systems. Also, like my peers across academia, I care deeply about the health of the peer review system.
I recently led an exploratory study that asked six OpenAI models to review 137 anonymized scientific manuscripts. These manuscripts were taken from the 2017 Annual Meeting of the Association for Computational Linguistics, or ACL.
Each model was prompted to review manuscripts one at a time, and the prompts were constructed using instructions taken directly from the ACL’s reviewer guidelines. For each manuscript, the models generated three independent reviews, including numeric scores ranging from 1 to 5 on nine criteria such as the overall recommendation, impact, and originality. This process was repeated for three different settings of the reasoning effort parameter—low, medium, and high—for the reasoning models (o1, o3-mini, o3, and GPT-5), and four different values of the ‘temperature’ parameter (which controls the amount of randomness in LLM output) for GPT-4o and GPT-4.1.
We didn’t give the models any other explicit constraints like a target acceptance rate or adopting a critical or friendly tone.
Yet, what we found was surprising. Four of the six models (GPT-4.1, GPT-4o, o1, and o3-mini) accepted nearly 100 percent of all papers. Two models (o3 and GPT-5) accepted nearly two-thirds of the papers.
This absurdly high acceptance rate raises serious doubts about the trustworthiness of AI for peer review, and does not bode well for science. It stands in stark contrast with the actual acceptance rate of less than 25 percent at the conference from where the manuscripts in the study were sourced.
In 2025, the AAAI, faced by the deluge of submissions, began a pilot program of incorporating LLMs in their peer review pipeline with an aim to augment the expert judgment of human peer reviewers.
The AAAI system used OpenAI’s GPT-5 embedded within a multilayer, automated AI review pipeline. After removal of submissions that were non-compliant with submission policies (were not anonymized, or exceeded the maximum manuscript length, for example) each of the remaining 22,977 papers received at least two human reviews and one review generated via the AI pipeline.
A recent arXiv preprint reports that AAAI’s system generated all reviews in less than 24 hours. A survey of the authors, committee members, and area chairs of the conference showed that they found the AI reviews useful and preferable to human reviews especially in their assessment of technical accuracy and in providing suggestions for improving the manuscripts.
About one in seven respondents from the program committee and chairs reported that the AI reviews changed their evaluation of the manuscript, whereas more than half said it did not. About half of this group also agreed that the AI reviews found issues that human reviewers would not have. As a counterpoint, the broader respondent pool suggested that the AI reviews displayed a propensity to overemphasize relatively minor issues.
Whereas the AAAI pilot represents a step forward in pointing out the pros and cons of integrating AI at the conference organizers’ end, it says nothing of the human reviewers and their potential use of AI to review the manuscripts. And while the AAAI’s efforts are commendable—and born of necessity—the preprint report doesn’t say much about the intrinsic values or preferences that shape the behaviors of the LLMs at the core of their AI review system.
LLMs are AI models, typically trained on billions of pieces of text spanning from digitized versions of Homer’s Odyssey, to computer-generated documents and reports, to your local ice cream shop’s blog posts.
After training, the LLMs are fine-tuned through preference judgments from human annotators in a process called Reinforcement Learning from Human Feedback, or RLHF. In this process, human annotators indicate which of several model responses to typical prompts they prefer and which not, based on criteria such as accuracy and safety.
These preference judgments, shaped together by the visions and values of the model vendors as well as those of the annotators, get baked into the models before they get shipped out to users.
The OpenAI models we examined in our study exhibited a strong emphasis on values that favor caring and tolerance for others’ views, benevolence, and self-directed thinking. These values seem to be aligned with AI safety principles such as respect and inclusive attitudes toward each other’s viewpoints—and rightly so.
However, peer review is grounded in values and ethical norms such as integrity, unbiased critique, and confidentiality. As more evidence points to absurdly high paper acceptance rates from LLMs, it raises the question whether the current crop of AI models operate on the same values as those in which peer review is grounded.
If these values are not shared or aligned, can AI be trusted for peer review?
The analysis of open-ended survey responses from the AAAI pilot revealed other concerns that should give us pause. First, authors could tailor their manuscripts so as to be assessed more favorably by AI systems, thus undermining the credibility of an already strained peer review system. The efficacy of such attempts at manipulating LLM reviewers has already been demonstrated through injection of imperceptible text and steering of LLM judgments through hidden instructions in manuscripts.
The second, and perhaps more important reason concerns novelty in research. The AAAI survey revealed a persistent concern expressed by respondents: that the AI reviews did a poor job of judging the novelty of the papers. LLMs have also been shown to disagree with human judgments of novelty in research grant proposals. This should ring alarm bells not only for peer reviewers, conference chairs, or journal editors, but also for grantmaking bodies that may be considering AI use for culling and evaluating proposals.
This begs a question that every person, scientist or not, must confront: if the pursuit of novelty is what propels the uniquely human scientific endeavor, then should we outsource its evaluation to machines?
As the key stakeholders, scientists, authors, peer reviewers, publishers, university administrators, grantmaking institutions—and perhaps those in the legal professions—need to weigh in on the questions raised by this research. Together we need to create robust guidelines, and ultimately sound policies for AI-driven automation which has serious implications for public trust in science and other important societal institutions.













Leave a Reply