Sitemap

My Paper Was Outdated Before It Was Published: Rethinking Peer Review

9 min readAug 5, 2025

--

Disclaimer: My peer review experience is mostly limited to the areas of statistics and machine learning. Other areas may have very different processes and expectations, and the discussion I make here does not apply to all fields.

“We submitted in March. Got the first review 8 months later. It was a desk reject.”

“A reviewer asked us to compare with a model that didn’t even exist when we submitted.”

“They requested three rounds of revision. By the time it was accepted, the preprint had over 50 citations.”

“A reviewer said our work wasn’t novel. That same reviewer later published a paper using our exact method.”

“We got a ‘major revision’ with no actual comments, just a request to cite the reviewer’s own papers.”

“After 20 months, we got a rejection because the topic was ‘no longer of interest.’”

“All reviewers recommended acceptance after three rounds (and two years!). Then the editor suddenly rejected it, coming up with new arguments.”

Every researcher has a version of these stories.

If you work in fast-moving fields like statistics or machine learning, this kind of experience isn’t an exception: it’s the norm. These delays aren’t just frustrating; they raise serious questions about whether traditional peer review still serves its intended purpose. In this post, I argue that the current system no longer aligns with the pace of modern science, and that we need to reconsider how we validate and disseminate research.

How Peer Review Started

Originally, peer review was introduced due to practical needs. When the first scientific journals were launched in the 17th century — like the Journal des Sçavans (Paris, 1665) and Philosophical Transactions of the Royal Society (London, 1665) — printing was expensive and limited. Editors couldn’t publish everything they received. To decide which manuscripts were worth printing, they consulted other scholars to assess the quality, novelty, and importance of submissions. This early practice wasn’t a formal peer review, but it served as a filter to manage the high costs of printing and the scarcity of journal pages.

Over time, peer review evolved. By the mid-20th century, especially after World War II, there was a huge increase in the number of researchers and in government funding for science. With more scientists submitting more papers, journals introduced a more formal and standardized peer review process. Anonymous external reviewers became the norm in the 1950s and 1960s.

This shift also changed the role of publication. It was no longer just a filter for what got printed. Instead, it became the main way to measure scientific merit. Universities and funding agencies needed easy ways to assess researchers, and publishing in prestigious journals became the most accepted indicator of success. A process that started as a practical necessity turned into a system that determined careers.

Do We Still Need Peer Review?

Today, we no longer have a shortage of publication space. Preprints and digital platforms allow anyone to publish their work instantly. Yet, the old review system remains, and its delays are especially painful in fast-moving fields. We all have many stories of papers that took years to get through the review process — by then, their models, data, or findings may already be outdated or overtaken by new developments.

As a result, many of the most influential papers are already widely read, discussed, and cited while still in preprint form on platforms like arXiv. By the time a paper is officially published, its preprint has often been cited and built upon. Discussions happen on social media, code is shared on GitHub, and researchers often reach a consensus on the value of a paper long before it appears in a journal. By the time of formal publication, the process may feel like little more than a bureaucratic rubber stamp. It is still useful for your CV, but almost irrelevant to real-time scientific progress.

Does Peer Review Improve Science?

The primary defense of peer review today isn’t about managing costs, but about ensuring scientific quality. But when you look at the data, the picture is not so clear.

  • It Does Catch Small Mistakes: Studies show peer review is fairly effective at improving a paper’s clarity, fixing typos, and cleaning up figures. But these are often superficial changes that don’t affect the core scientific claims of the paper (Baxt & Waeckerle, 1994). But in an age where large language models can perform these surface-level checks in seconds, using a human expert’s time for this is a waste.
  • It Often Fails to Detect Big Errors: Controlled experiments indicate that reviewers often fail to identify serious flaws. For instance, in one experiment, researchers at the BMJ sent a paper with nine major errors to a group of reviewers. On average, reviewers only caught 2.58 of the nine flaws (Godlee et al., 2006). This suggests peer review is not the reliable safety net we believe it to be.
  • It’s Highly Inconsistent: One of the biggest issues is how random the process is. Studies have found that the agreement between two reviewers looking at the same paper is often barely better than a coin toss. The kappa value (a measure of inter-rater reliability) is often around 0.17 (Bornmann et al., 2010), indicating very little agreement. Whether a paper gets accepted can often come down to the “luck of the draw”, that is, who happens to be assigned as your reviewer.

A famous experiment from the 2014 NeurIPS conference (a top machine learning conference with very low acceptance rates) illustrates this randomness (Lawrence, 2022). The organizers had two independent committees review the same batch of papers. The results were worrying:

  • For over a quarter of the papers, the two committees came to opposite conclusions (one accepted, the other rejected).
  • If you took a paper that was accepted by a committee, there was a 49.5% chance it would have been rejected if it had been reviewed by the other committee instead.

To make things worse, there’s evidence that truly innovative work is more likely to be rejected by this consensus-driven system (Fang et al., 2020).

The Root of the Crisis

The crisis in peer review comes from two main sources.

First, even with the best intentions, reviewing is inherently subjective. No one truly knows which ideas will become foundational for future discoveries. Many papers that were initially rejected or dismissed as unimportant later became landmarks in their fields. When Peter Higgs submitted his 1964 paper outlining the “Higgs mechanism”, it was rejected by Physics Letters, with an editor reportedly deeming it of “no obvious relevance to physics” (https://www.ph.ed.ac.uk/higgs/brief-history). Nearly 50 years later, after the confirmation of the Higgs boson at CERN, he was awarded the Nobel Prize.

But the issues run deeper than subjectivity. The quality of reviews is widely perceived to be declining in machine learning and statistics. Editors often say it’s getting harder to find good reviewers. Why? Because there’s no real incentive. Researchers are judged primarily on their publications, not their reviews. Reviewing is unpaid, anonymous, and essentially thankless work that takes valuable time away from the very research needed to advance a career.

The problem is worsening due to the sharp rise in submissions. As the figures for arXiv and NeurIPS submissions show, the volume of papers has exploded, while the number of qualified experts has not increased at the same rate.

Press enter or click to view image in full size
Monthly Number of arXiv Submissions. Data Source: https://arxiv.org/stats/monthly_submissions
Press enter or click to view image in full size
Yearly Number of NeurIPS Submissions. Data Source: https://papercopilot.com/statistics/neurips-statistics/neurips-2024-statistics/

The consequence is often rushed, low-quality reviews that make the “luck of the draw” problem even worse. Finally, we can’t ignore the instances of bad faith: reviewers who hold up a competitor’s paper, demand citations to their own (often unrelated) work, or provide unconstructive, dismissive feedback.

Alternatives to the Traditional Model

We don’t need to rely on peer review to do science. Several platforms are already experimenting with faster, more open, and more dynamic models. Here are some examples:

  • arXiv: The original preprint server, it forms the foundation of open access. It makes science accessible in real-time.
  • SciRate: Platforms like SciRate build directly on arXiv, adding a community-based layer of discussion and rating that helps to separate the signal from the noise without formal gatekeepers.
  • OpenReview: This is the platform that powers the shift toward transparency for many top ML conferences and new journals. Instead of a secret process, OpenReview hosts the entire review life cycle publicly. Submissions, reviews, author rebuttals, and final decisions are open for anyone to read. While some problems persist (such as reviewer assignment variability), depending on how the platform is used, it represents a meaningful improvement over the traditional peer review model.
  • F1000Research & eLife: These platforms are pioneering a “publish, then review” model. Manuscripts are posted publicly after a basic screening, and then invited, signed reviews are published alongside the article. This transparency not only improves review quality but also gives reviewers formal credit for their work.

So, Where Do We Go From Here?

If better alternatives are already here, why are we stuck in the old system? The answer is a mix of inertia and incentives. Concrete academic success, such as hiring, tenure, and funding, is still mostly measured by publications in prestigious journals that rely on traditional peer review.

This creates a dilemma, especially for early-career researchers. You can see that the system is in need of redesign, but you can’t afford to opt out. However, promising middle-ground options are starting to get attention. A key example is Transactions on Machine Learning Research (TMLR), which seems to be getting traction. TMLR uses OpenReview (which reduces the chances of low quality reviews), and its primary acceptance criterion is whether a paper’s claims are correct and well-supported, not a subjective judgment of its potential impact. This directly addresses some of the problems I discussed above. Additionally, all reviews and discussions are published alongside accepted papers, providing full context for the research.

Despite such progress, the broader system still fosters a major disconnect: the tools we use to do science are advancing at light speed, while the system we use to validate science remains mostly stuck in the past. We’re spending our time, energy, and public funding on a process that is often unreliable and controlled by publishers who often reap profits from our free labor (a subject for a separate post). Therefore, the goal isn’t to abolish peer review, but to reimagine it.

We need to move from a mindset of gatekeeping to one of ongoing conversation. The future should be built on a foundation of transparency and collaboration, where review is a continuous, public dialogue, not a one-time, anonymous verdict. We have the tools to make this happen, but technology alone isn’t enough. It requires a shift in how we, as a community, define and reward scientific contribution.

How do we break this cycle? What concrete steps can universities, funding bodies, and individual researchers take to build a review system that actually accelerates discovery instead of hindering it? I’m curious to hear your thoughts.

References and Additional Sources

1. Goodman, S. N., Berlin, J. A., Fletcher, S. W., & Fletcher, R. H. (1994, July 1). Manuscript quality before and after peer review and editing at Annals of Internal Medicine. Annals of Internal Medicine, 121(1), 11–21. https://doi.org/10.7326/0003‑4819‑121‑1‑199407010‑00003 PubMed

2. Smith, R. (2006, April). Peer review: A flawed process at the heart of science and journals. Journal of the Royal Society of Medicine, 99(4), 178–182. https://doi.org/10.1177/014107680609900414 PubMed

3. Bornmann, L., Mutz, R., & Daniel, H.‑D. (2010, December). A reliability‑generalization study of journal peer reviews: A multilevel meta‑analysis of inter‑rater reliability and its determinants. PLOS ONE, 5(12), e14331. https://doi.org/10.1371/journal.pone.0014331 PLOS

4. Lawrence, N. D. (2022, May 10). The NeurIPS Experiment (SNSF). Inverse Probability. Retrieved August 3, 2025, from https://inverseprobability.com/talks/notes/the-neurips-experiment-snsf.html

5. Brezis, E. S., & Birukou, A. (2020, April). Arbitrariness in the peer review process. Scientometrics, 123(1), 393–411. https://doi.org/10.1007/s11192-020-03348-1 link.springer.com

6. Wasserman, L. (2012, February 20). A World Without Referees [essay]. Carnegie Mellon University. Retrieved August 3, 2025, from https://www.stat.cmu.edu/~brian/zach-00/765-2019/764-2016/week05/refereeing/Larry-Abolish-the-Peer-Review.pdf

https://x.com/authorea/status/1086372988159774720/photo/1

--

--

Rafael Izbicki
Rafael Izbicki

Written by Rafael Izbicki

Associate Professor at UFSCar, PhD from CMU, CNPq Research Fellow. I work on theory, methods, and applications in statistics, ML, and data science.