Open Access Review
Review of “Scientific Paper”
computer science
Original paper: https://arxiv.org/pdf/2605.07723
Overall assessment
Overall, this is a potentially high-impact paper with a compelling empirical contribution, but it would benefit from substantial revision to improve methodological clarity, statistical reporting, and the careful framing of claims beyond the studied corpora.
Strengths
The paper has a strong, timely premise and a clear contribution: it tackles an important real-world problem with a large-scale, cross-corpus audit of citation hallucinations, and the Introduction and Discussion present the motivation, novelty, and broader significance especially well.
Points to improve
The main weaknesses are in methods transparency and results reporting: the verification pipeline, statistical modeling, and uncertainty quantification need fuller specification, and the Results section should separate descriptive findings from interpretation more cleanly.
Recommendations
- State the review type explicitly and align it with the actual objective; if this is intended as a review, provide a protocol or registration identifier, and describe the study identification, screening, and inclusion criteria in full. Also add a formal risk-of-bias or quality appraisal plan and explain how it will inform synthesis.
- State the specific statistical test or model used for each main comparison and link each result sentence to that analysis; report exact sample sizes, denominators, and uncertainty measures for key estimates; and specify any multiple-comparison corrections, matching criteria, or model covariates that materially affect the reported results.
- Tighten the interpretation in the Discussion by explicitly labeling cross-domain implications as speculative, linking each major implication more directly to the specific result that supports it, and reducing certainty when discussing systemic effects beyond the studied corpora.
- Add a concise, study-specific generalizability statement and clarify the limitations of the reference-verification pipeline: explain that the title-based design could undercount hallucinations in underindexed fields, niche venues, and mathematical literature; clarify the boundary between unverifiable titles and real citations that misrepresent source content; and state that the current study captures only part of the broader citation-hallucination problem.
- Explicitly link each future direction to a specific empirical result or limitation from the study, and for each proposed follow-up specify the study design, target setting, and key outcome measures more concretely; separate near-term actionable directions from longer-horizon questions to make the future-research agenda easier to follow.
Paper summary
This study audits the real-world prevalence of LLM hallucinations by focusing on scientific citations, an object with verifiable existence. The authors analyze 111 million references from 2.5 million papers across arXiv, bioRxiv, SSRN, and PubMed Central, using a multi-step verification pipeline that compares extracted citations against bibliographic databases and Google Scholar while correcting for baseline matching errors seen before widespread LLM use. The analysis shows a sharp increase in non-existent references beginning in 2023 and especially in mid-2024, with hallucinated citations appearing diffusely across many papers rather than being confined to a few highly contaminated ones. The phenomenon is more common in fields with higher inferred LLM use, among small and early-career author teams, and in papers with linguistic markers of AI-assisted writing. The study also examines who receives credit from these hallucinations, finding that invented or misattributed citations disproportionately favor already prominent, highly cited, and male scholars. It further evaluates safeguards, showing that moderation and peer review catch only part of the problem and that many hallucinated citations persist from preprint to publication and even enter bibliographic databases.
Main claims
- LLM adoption is associated with a sharp rise in non-existent scientific references across major preprint and publication corpora.
- The study estimates at least 146,932 hallucinated citations in 2025 alone across the analyzed datasets.
- Hallucinated citations are diffusely distributed across many papers rather than concentrated in a few severely contaminated manuscripts.
- Citation hallucinations are more common in fields with higher inferred LLM use, in papers with linguistic signs of AI-assisted writing, and in smaller or early-career author teams.
- Hallucinated references disproportionately credit prominent and male scholars, suggesting that LLM errors may reinforce existing inequities in scientific recognition.
- Existing moderation and peer review systems detect only a fraction of hallucinated citations, indicating that current safeguards are insufficient.