Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
Jaeho Kim (UNIST), Yunseok Lee, Seulki Lee
Overview
In this insightful and timely talk at ICML 2025, Jaeho Kim, along with Yunseok Lee and Seulki Lee from UNIST South Korea, presented a compelling position paper addressing the escalating crisis in AI conference peer review. Titled "The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards," the presentation critically dissects the systemic failures that increasingly render paper acceptance dependent on a "reviewer lottery" rather than the intrinsic quality of research. The core argument posits that without immediate and significant structural reforms, the integrity and fairness of top-tier AI conferences are at risk, particularly impacting junior researchers.

Key moments
- 0:00 Introduction: The AI conference reviewer lottery crisis
- 2:00 Identifying root causes of the peer review crisis
- 4:00 Proposal 1: Bidirectional author feedback system
- 6:00 Proposal 2: Digital badges and reviewer impact scores
- 6:27 Addressing alternative views and objections to proposals
- 8:00 Q&A: Discussing financial stipends for reviewers
Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
Speakers: Jaeho Kim, Yunseok Lee, Seulki Lee
Conference: ICML 2025
YouTube: https://slideslive.com/39043884
Overview
In this insightful and timely talk at ICML 2025, Jaeho Kim, along with Yunseok Lee and Seulki Lee from UNIST South Korea, presented a compelling position paper addressing the escalating crisis in AI conference peer review. Titled "The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards," the presentation critically dissects the systemic failures that increasingly render paper acceptance dependent on a "reviewer lottery" rather than the intrinsic quality of research. The core argument posits that without immediate and significant structural reforms, the integrity and fairness of top-tier AI conferences are at risk, particularly impacting junior researchers.
The speakers meticulously outline the multifaceted nature of this crisis, attributing responsibility to authors, reviewers, and the conferences themselves. While acknowledging the role of authors in the sheer volume of submissions, their primary focus—and the proposed solutions—target the profound power imbalance experienced by reviewers and the systemic lack of accountability and incentives within the review process. To counteract these issues, Kim and his colleagues propose a bidirectional feedback system, empowering authors to evaluate review quality, and a systematic reviewer reward system, designed to foster accountability and motivate higher quality contributions from reviewers. This talk is a vital call to action for the AI research community to collectively address and rectify the fundamental mechanisms governing scientific dissemination and recognition.
Background
▶ Watch: Introduction: The AI conference reviewer lottery crisis (0:00)
The landscape of AI/ML research has undergone an explosive transformation in recent years, marked by an unprecedented surge in both the quantity and complexity of submitted papers to premier conferences. This talk highlights a critical consequence: the current peer review system, designed for a different era, is now buckling under the immense pressure. The problem is not merely one of scale; it's a systemic breakdown where the foundational pillars of fair and rigorous evaluation are eroding.
Jaeho Kim opens with a stark, satirical vision of career advice for junior researchers in the year 2026, painting a dystopian picture where "patience is key" because papers must be submitted "four to five times" to overcome a mere 20% acceptance ratio, implying a dependency on luck rather than merit. He even humorously (and chillingly) suggests embedding "really small and white" text in papers for AIs to read, a direct jab at the growing concern over superficial or LLM-generated reviews. This opening serves to underscore the gravity of the crisis, which, if left unaddressed, threatens to undermine the very credibility of scientific peer review.
The speakers identify three primary culprits in this crisis:
- Authors: The sheer volume of submissions has "skyrocketed," a trend that has not been met with a proportional increase in the number of qualified reviewers. While this aspect is acknowledged, the paper primarily focuses on systemic solutions for the other two.
- Reviewers: A "strong power imbalance" exists, where reviewers face "no accountability for their actions." This unchecked power enables some to submit "superficial reviews" or, more recently, to "simply rely on LLM-generated reviews" without fear of repercussion. The lack of proper acknowledgment or reward further exacerbates this issue, disincentivizing diligent effort.
- Conferences: The systems currently in place neither effectively hold reviewers accountable for their quality of work nor adequately reward their efforts. This systemic oversight creates a vacuum where poor review practices can flourish unchecked, contributing to the "reviewer lottery" phenomenon.
Prior work in improving peer review has often focused on optimizing assignment algorithms or reviewer training. However, this position paper argues for a more fundamental shift, addressing the core issues of accountability and incentives, which have largely remained unaddressed in a comprehensive, systematic manner within the AI/ML conference ecosystem. The problem, therefore, is not just about finding enough reviewers, but about ensuring that the reviews themselves meet a high standard of quality, fairness, and intellectual rigor.
Key Findings
▶ Watch: Proposal 1: Bidirectional author feedback system (4:00)
The central finding of this position paper is the diagnosis of a full-blown "AI Conference Peer Review Crisis," characterized by a dangerous shift where paper acceptance is becoming "more dependable on reviewer lottery than research quality." This crisis is multifaceted, stemming from the exponential growth in submissions, a stagnating pool of qualified reviewers, and critical systemic flaws within the review process itself.
The speakers identify the following key contributing factors and their consequences:
- Unsustainable Growth: The number of submissions to top-tier AI conferences has exploded (e.g., NeurIPS receiving 25,000 submissions), far outstripping the capacity of the existing pool of qualified reviewers. This imbalance strains the system and compromises review quality.
- Reviewer Accountability Deficit: A significant "power imbalance" exists between reviewers and authors. Reviewers currently face "no accountability for their actions," leading to irresponsible behaviors such as submitting superficial reviews or relying on LLM-generated content without proper scrutiny. This lack of oversight directly contributes to the "reviewer lottery" effect.
- Lack of Reviewer Incentives: The current system fails to adequately acknowledge or reward reviewers for their often-extensive and critical work. This absence of visible, academically valuable recognition demotivates reviewers and discourages the investment of time and effort required for high-quality evaluations.
- Absence of Author Safeguards: Authors currently possess minimal to no mechanisms to challenge or flag irresponsible or low-quality reviews effectively. This leaves authors vulnerable to unfair evaluations and inhibits their ability to ensure their work is properly understood.
- Evidence Gap: While anecdotal evidence and social media discussions suggest a decline in review quality, the speakers acknowledge the absence of concrete, qualitative evidence. However, they equally stress that there is "no evidence that the quality of the review is being maintained," calling for comprehensive surveys by conference organizers to quantify the problem.
In response to these critical findings, the paper proposes two interconnected and innovative solutions:
- A Bidirectional Feedback System: This system aims to introduce accountability by allowing authors to evaluate the quality of the reviews they receive, providing a much-needed safety measure against irresponsible or LLM-generated content.
- A Systematic Reviewer Reward System: This proposal seeks to motivate reviewers through visible and academically valuable recognition, such as digital badges and reviewer impact scores, thereby fostering a culture of higher quality and engagement.
These findings collectively highlight a pressing need for structural reform to uphold the scientific integrity and fairness of the AI/ML research ecosystem.
Technical Deep Dive
▶ Watch: Proposal 2: Digital badges and reviewer impact scores (6:00)
The proposed solutions are designed to systematically address the identified failures in the AI conference peer review process through two primary mechanisms: a bidirectional feedback system and a systematic reviewer reward system. While this is a position paper, the "technical deep dive" focuses on the design and operational flow of these proposed systems, rather than specific model architectures or algorithms for a traditional ML system.
Bidirectional Feedback System
The current peer review model follows a linear path: paper submission, reviewer assignment, review generation, simultaneous release of all reviews, author-reviewer discussion, and a final decision. The proposed bidirectional feedback system introduces critical intervention points and a novel approach to review disclosure:
- Incorporating LLM-generated Reviews (Paradoxically): A core tenet of the proposal is to "paradoxically incorporate LLM-generated reviews alongside human reviews." The rationale is that LLM-generated reviews, in their current state, are often "very superficial." By providing an official, conference-generated LLM review as a reference, authors gain a baseline. If a human reviewer's submission is also superficial or strikingly similar to the LLM-generated reference, authors can "have more systematical measures to challenge this author [reviewer] and also report to the conference." This mechanism aims to empower authors to identify and flag irresponsible, low-effort reviews, especially those potentially outsourced to AI without human oversight.
- Split Release of Reviews: Instead of a single, simultaneous release of all review components, the system proposes a three-stage disclosure process:
- First Release: Only the "very neutral parts of the reviews" are disclosed to authors. This includes the summary of the paper, its identified strengths, and any clarifying questions from the reviewer. Notably, weaknesses and initial ratings are withheld at this stage.
- Author Feedback and Flagging: Following the first release, authors are given a window to "rate the reviews that they have received and flag any irresponsible reviews." This is the crucial point where authors can provide feedback on review quality, potentially identifying superficiality, miscomprehension, or suspected LLM-generated content using the official LLM reference review.
- Second Release: Only after author feedback has been submitted, the "rest of the reviews such as the weaknesses and ratings are promptly disclosed." This staggered approach ensures that author feedback is collected on the quality of comprehension and engagement rather than solely on the subjective rating or criticism of their paper.
The rational behind this design is multifold:
- Minimal Safety Measure: It creates a much-needed "minimal safety measure to protect authors from irresponsible reviews," which is currently absent.
- Evaluation of Comprehension: Authors are uniquely positioned to "evaluate whether the reviewers have properly comprehended their paper." By focusing on summaries, strengths, and questions initially, authors can assess if the reviewer genuinely engaged with the content.
- Reference for Challenge: The official LLM-generated review provides a concrete "reference to challenge LLM-generated reviews" from human reviewers.
The system is envisioned to be beneficial in both the short and long term: immediately filtering out irresponsible reviews and, with sufficient scale, enabling Associate Chairs (ACs) and Program Chairs (PCs) to "make better decisions on papers" by incorporating review quality into their final assessments.
Systematic Reviewer Reward System
The second major proposal focuses on motivating reviewers through enhanced recognition and academic value, moving beyond the current system where reviewer contributions are largely invisible.
- Digital Badges:
- Concept: Digital badges would be awarded to reviewers to acknowledge their service.
- Visibility: These badges would be "displayed alongside the open review and Google Scholar profiles." This ensures public recognition and academic visibility for review contributions.
- Impact: Research on digital badges suggests that such systems lead to "enhanced participation and an enhanced sense of community," which is precisely what the peer review process requires. They serve as a low-cost, high-impact motivational tool.
- Reviewer Tracking and Reviewer Impact Scores:
- Concept: This proposes a "long-term benefit to track reviewer activities" and provide "quantitative measures for reviewer services."
- Analogy to H-index: The speakers explicitly draw an analogy to the H-index, a widely accepted metric for author impact based on citations. For reviewers, the proposed "reviewer impact score" would increase "if the paper we have reviewed gets more citation." This directly links a reviewer's impact to the subsequent success and influence of the papers they rigorously evaluated.
- Academic Value: By creating a quantitative, academically recognized metric for reviewing, this system aims to elevate the status of reviewing from a thankless task to a valuable, career-enhancing contribution.
Addressing Alternative Views: The speakers proactively address common counterarguments:
- Conferences handling 10,000+ submissions: They argue that handling a large volume doesn't equate to having a proportional number of qualified reviewers.
- No evidence of declining review quality: They counter that there's also "no evidence that the quality of the review is being maintained," urging conference organizers to conduct comprehensive surveys.
- Fear of difficult reviewer recruiting: While acknowledging that implementation "could initially make reviewer recruiting more difficult," they assert that "in the long term it will be helpful to the community" by professionalizing the role and offering tangible rewards.
Notably, the speakers reject the idea of stipends for reviewers, citing concerns about conference budgets and the potential burden on authors from less resourced countries or labs if costs were passed on. This pragmatic stance underscores their focus on sustainable, system-level changes that don't introduce new financial barriers.
In essence, the technical deep dive reveals a well-thought-out systemic overhaul that integrates feedback, accountability, and recognition to professionalize and elevate the AI conference peer review process.
Experimental Setup & Results
▶ Watch: Addressing alternative views and objections to proposals (6:27)
As a position paper, this talk does not present empirical experimental setups or results in the conventional machine learning research sense. Instead, it critically analyzes the current state of AI conference peer review and proposes conceptual systemic changes based on logical reasoning, observed trends, and insights into human motivation and behavior.
The "results" of this work are therefore theoretical and propositional: the identification of a crisis, the articulation of its root causes, and the detailed design of two proposed solutions—the bidirectional feedback system and the systematic reviewer reward system. The talk serves as a call to action and a blueprint for future implementation and empirical validation by conference organizers.
The authors acknowledge the lack of "concrete or qualitative evidence that the review quality is declining" but also emphasize the converse: "there is also no evidence that the quality of the review is being maintained." This highlights the need for a comprehensive survey targeting both authors and reviewers to gather the necessary data to empirically validate the extent of the crisis and, subsequently, the effectiveness of any proposed solutions.
While no direct experiments were conducted, the presentation draws on broader research for its claims, such as the finding that "Research on digital badges show that such system leads to enhanced participation and an enhanced sense of community," supporting the rationale for their inclusion in the reward system. The talk's impact lies in its strong argument for a paradigm shift, prompting the community to consider and potentially implement these structural changes, whose "results" would then be observed in the improved quality and fairness of future peer review cycles.
Practical Implications
▶ Watch: Q&A: Discussing financial stipends for reviewers (8:00)
The proposals outlined in this position paper carry significant practical implications for all stakeholders in the AI research ecosystem, from individual researchers to large-scale conference organizers. The overarching goal is to professionalize and elevate the peer review process, making it more equitable, transparent, and ultimately, more effective at identifying and disseminating high-quality research.
For Practitioners (Authors)
- Reduced Reviewer Lottery: Authors could experience a fairer review process, with less dependence on the luck of reviewer assignment. The bidirectional feedback system aims to filter out superficial or irresponsible reviews, ensuring that papers are judged on their merits.
- Mechanism for Challenge: The proposed system provides authors with a concrete, systematic way to challenge low-quality reviews, especially those suspected of being LLM-generated or demonstrating a clear lack of comprehension. This empowerment reduces the current vulnerability of authors to unfair evaluations.
- Improved Feedback Quality: With increased reviewer accountability and incentives, authors can expect to receive more constructive, detailed, and insightful feedback, which is crucial for improving their work.
For Infra Teams (Conference Organizers and Platforms like OpenReview)
- System Development & Integration: Implementing the proposed systems would require significant development effort. Conference platforms would need to build new interfaces for author feedback, integrate LLM-generated review capabilities, manage staged review disclosures, and develop systems for tracking reviewer activity and awarding digital badges and impact scores.
- Increased Operational Overhead (Initially): While aiming for long-term efficiency, the initial deployment and management of these new systems might introduce additional operational complexity and resource demands.
- Data Collection & Analysis: Organizers would need to conduct the comprehensive surveys advocated by the speakers to establish baselines for review quality and continuously monitor the impact of the new systems.
- Policy Development: Clear policies would be needed regarding how author feedback on reviews is used (e.g., influencing AC decisions, reviewer warnings/bans) and the precise methodology for calculating reviewer impact scores.
For Model Builders (Reviewers)
- Enhanced Accountability: Reviewers would face greater scrutiny and accountability for the quality of their work. This is intended to motivate higher standards and discourage superficial or irresponsible practices.
- Tangible Recognition & Motivation: The introduction of digital badges and reviewer impact scores provides concrete, visible, and academically valuable recognition for diligent review service. This can serve as a powerful motivator, transforming reviewing from a thankless task into a recognized contribution to the scientific community, potentially aiding career progression.
- Professional Development: The expectation of higher quality reviews, coupled with feedback mechanisms, could encourage reviewers to invest more in their review skills, leading to personal and professional growth.
Tradeoffs and Limitations
- Initial Reviewer Recruitment Challenges: The speakers acknowledge that implementing these systems could "initially make reviewer recruiting more difficult" as some may resist increased accountability. However, they posit that it will be beneficial in the long term by professionalizing the role.
- Implementation Complexity & Cost: Developing and maintaining these new functionalities (e.g., LLM review generation, author feedback UI, badge/score tracking) represents a non-trivial investment for conference organizers, many of whom, as noted in the Q&A, are "underfunded and a bit understaffed."
- Potential for Misuse: While "safety measures for both authors and reviewers" are mentioned, the specifics of preventing authors from unfairly flagging reviews, or reviewers from gaming the impact score system, would need careful design.
- Defining Reviewer Impact: While a citation-based H-index analogy is proposed, the nuances of reviewer impact are complex. A highly critical review that leads to a significant revision might be more impactful than a simple acceptance of a highly cited paper. The metrics would need to evolve to capture this complexity.
- LLM Review Quality: The effectiveness of using LLM-generated reviews as a reference relies on their consistent "superficiality." As LLMs advance, their ability to generate convincing, albeit unoriginal, reviews might complicate this mechanism.
Despite these challenges, the practical implications point towards a future where AI conference peer review is more robust, fair, and transparent, ultimately bolstering the integrity and quality of AI research.
Key Takeaways
- The AI conference peer review system is facing a significant crisis, primarily driven by an exponential increase in submissions, a lack of qualified reviewers, and critical systemic issues related to reviewer accountability and incentives.
- The current system often leads to paper acceptance being determined by a "reviewer lottery" rather than by intrinsic research quality, posing a threat to the integrity of top-tier AI conferences.
- A bidirectional feedback system is proposed, empowering authors to evaluate review quality. This system includes a staged release of review components (neutral parts first, then weaknesses/ratings after author feedback) and the paradoxical incorporation of LLM-generated reviews as a reference for authors to identify superficial or AI-generated human reviews.
- A systematic reviewer reward system is advocated to motivate higher quality reviews. Key components include digital badges for visible recognition (displayed on open review platforms and Google Scholar profiles) and reviewer impact scores, analogous to an H-index, which would increase based on the citation count of papers reviewed.
- While acknowledging that implementing these changes might initially make reviewer recruitment more challenging and require significant effort from conference organizers, the long-term benefits of enhanced accountability, improved review quality, and a stronger sense of community within the AI research ecosystem are deemed essential and outweigh initial hurdles.
- The speakers urge conference organizers to conduct comprehensive surveys targeting both authors and reviewers to gather empirical data on current peer review quality, underscoring the need for evidence-based reform.
About the Speaker(s)
The talk was presented by Jaeho Kim from UNIST South Korea, who served as the primary speaker. He was joined by co-authors Yunseok Lee and Seulki Lee, also affiliated with UNIST. Their collaborative work focuses on critically analyzing and proposing solutions to the pressing issues within the AI conference peer review process. Through their position paper, they advocate for significant structural reforms to improve fairness, accountability, and quality in the scientific dissemination of AI/ML research.
Reviews
Maya Iyer (Theoretical ML Researcher) — WEAK
A position paper diagnosing real dysfunction in ML conference peer review and proposing two structural remedies — bidirectional author feedback and a reviewer reward system. The diagnosis is widely shared and the motivation is genuine, but the proposed solutions are underspecified, the supporting evidence is thin to nonexistent, and the core mechanisms introduce new failure modes that receive only cursory treatment. This is advocacy dressed as analysis.
Chen Zhao (Applied ML Researcher & Empiricist) — WEAK
A position paper diagnosing real dysfunction in AI conference peer review and proposing two structural interventions — author feedback on reviews and a reviewer reward system — but with essentially no empirical grounding for either the diagnosis or the proposed remedies. The problem is real and the community conversation is worth having, but the work as presented is advocacy with citations, not a research contribution.
→ Top-rated talks at International Conference on Machine Learning 2025
All talks from International Conference on Machine Learning 2025