Open Access Review
Review of “Scientific Paper”
iot
Original paper: https://arxiv.org/pdf/2602.1961322255ggg
Overall assessment
With major revisions that address study-design limitations, improve statistical reporting and alignment with the stated research question, expand and sharpen the limitations and temper conclusions, the manuscript would make a useful, high-impact contribution to debates about LLMs in scholarly publishing; as currently presented it requires major revision.
Strengths
The manuscript asks an important, timely question and is well framed: the Introduction is clear, well referenced, and situates the study within orthopaedic publication ethics and recent LLM performance literature; the experimental procedure (blinded circulation of a hybrid manuscript) is straightforward and the results and data presentation (tables/figure) report medians, IQRs and reviewer decisions concisely.
Points to improve
However, the study design, analysis, and interpretation need strengthening — critical issues include an under-justified small reviewer sample and purposive sampling, an imbalance because the non-AI sections had prior peer review, sparse and potentially inappropriate statistical reporting (unnamed test, single p-value, no effect sizes or CIs), and some over-claiming in the Discussion that extends beyond what the data support.
Recommendations
- State the primary endpoint and how reviewer identification or acceptance will be measured.
- Specify the sampling frame, inclusion criteria, and rationale for selecting six reviewers.
- Align the statistical test with the ordinal reviewer scores and define any secondary analyses.
- Report the exact p-value along with the relevant sample size and, if available, an effect size or confidence interval.
- State explicitly how the prior peer review of the non-AI sections could have advantaged the human-written content and confounded comparisons with the AI-generated discussion. Explain why this design difference should temper conclusions about the quality of AI output.
Paper summary
This study evaluated whether an advanced large language model, ChatGPT-4, could generate a scientific discussion and conclusion for an orthopaedic article that would withstand blinded peer review in a high impact journal. The authors took the introduction, methods, and results from a recently published hip arthroplasty paper and asked ChatGPT-4 to produce a discussion and conclusion in the style of the Bone & Joint Journal. Six experienced fellowship trained arthroplasty surgeons then independently scored the five sections of the hybrid manuscript without knowing that part of it was AI generated. The AI written discussion and conclusion received somewhat lower median scores than the human written sections, but the difference was not statistically significant. Most reviewers recommended acceptance after major revisions, and none recommended outright rejection. The reviewers did not identify the AI generated sections as computer written. The paper argues that current AI models can generate plausible scientific prose that may pass experienced peer review, while also raising concerns about fabricated or inadequate references, limited comprehension of the study data, and dependence on the quality and recency of training data.
Main claims
- Current AI large language models can generate a discussion and conclusion that experienced blinded orthopaedic reviewers judge as suitable for publication after major revision in a high-impact journal.
- The AI-generated discussion and conclusion received lower median scores than the human-written sections, but the difference was not statistically significant.
- Blinded reviewers did not identify the AI-generated portions as computer generated and no reviewer recommended outright rejection.
- A major concern is that AI writing may rely on pattern recognition and may include fabricated, inadequate, or incomplete references, especially for novel topics.