1898 issues across 85 AI reviews: methodology and statistics dominate critical flags
Across 85 manuscripts, ManuscriptMind flagged 1,898 issues, with methodology and statistics carrying most of the critical load while writing problems were usually minor. Average rubric scores around 3.5 suggest drafts are mid-stage: promising but not yet ready for a demanding reviewer.
- 1,898
- Issues flagged
- 85
- Manuscripts reviewed
- 22.3
- Issues per manuscript
- 6.4%
- Rated critical
Across 85 manuscripts, ManuscriptMind identified 1,898 issues, with methodology alone accounting for 514 of them. The most consequential problems cluster in study design and analysis, while writing and literature gaps are common but rarely critical.
Figure 1
Issues by category and severity
Show data tableHide data table
| Category | Critical | Major | Minor | Total |
|---|---|---|---|---|
| Methodology | 90 | 364 | 60 | 514 |
| Statistics | 20 | 194 | 131 | 345 |
| Data presentation | 2 | 136 | 182 | 320 |
| Conclusions | 8 | 167 | 88 | 263 |
| Writing | 1 | 43 | 219 | 263 |
| Literature | 1 | 63 | 129 | 193 |
Where do the truly high-risk problems sit?
Figure 1 shows that methodology and statistics together carry much of the critical weight. Methodology issues are not just frequent, they are often severe, with a substantial share rated critical or major. This pattern is what you would expect when core design choices are still in flux. Problems like unclear inclusion criteria, inappropriate controls, or mismatched outcomes cannot be fixed by polishing prose. They threaten the validity of the results, so the tool consistently escalates them.
Statistical issues form a second cluster of high-stakes problems. Here the mix skews toward major rather than critical, suggesting that many analyses are structurally sound but under-specified or under-reported. Typical examples include missing power justifications, incomplete model descriptions, or ambiguous handling of missing data. These are the kinds of issues that an experienced reviewer will challenge before recommending acceptance, even if the underlying data are solid.
By contrast, categories like writing and data presentation are heavily populated with minor issues. That does not mean they are unimportant. Instead, it indicates that many manuscripts have a reasonably coherent narrative and figures that are interpretable, but with numerous small problems that accumulate: unclear variable labels, inconsistent terminology, or figure captions that do not fully describe the analysis. These rarely invalidate a study, yet they shape how much trust a reviewer places in the work.
Figure 2
Mean rubric score by domain
Show data tableHide data table
| Domain | Mean score (of 5) |
|---|---|
| Rigor | 3.52 |
| Significance | 3.71 |
| Presentation | 3.53 |
What does a mid-range rubric score actually mean?
The rubric averages sit in a narrow band: rigor at 3.52, presentation at 3.53, and significance at 3.71 on a 1 to 5 scale. This clustering is informative. It suggests that most manuscripts are neither deeply flawed nor close to journal-ready. Instead, they occupy a middle zone where the core idea has potential, but execution and exposition both need tightening.
A rigor score around 3.5 aligns with the severity mix in methodology and statistics. Many designs are conceptually appropriate but lack the level of detail and justification that peer reviewers expect. For authors, this is a signal to move beyond "standard methods" descriptions and explicitly justify choices like sample size, model selection, and outcome definitions.
The slightly higher significance score indicates that, on average, projects are asking questions that matter in their fields. The bottleneck is less about topic choice and more about demonstrating that the methods and analysis are strong enough to support the claimed contribution. When significance outruns rigor, reviewers are particularly likely to push for additional analyses or more cautious conclusions.
Presentation scores in the same mid-range suggest that most manuscripts are readable but not yet maximally transparent. Reviewers can follow the argument, yet still encounter friction in the form of dense paragraphs, unclear figure referencing, or insufficiently structured results sections. These are fixable problems, but they require deliberate revision rather than copyediting alone.
How should authors prioritize revisions before submission?
Taken together, the issue categories and rubric scores point to a clear revision strategy. First, treat methodology and statistics as the primary risk areas. Before worrying about stylistic polish, check that your design choices are explicit, justified, and aligned with your research question. If ManuscriptMind flagged a critical or major issue in these domains, that is the part of the manuscript most likely to trigger a negative reviewer response.
Second, use the abundance of minor issues in writing, data presentation, and literature to guide a structured clean-up. Multiple minor flags often indicate systemic patterns: for example, inconsistent reporting of units, incomplete legends, or citations that do not clearly support the statements they accompany. Addressing these in a batch can substantially improve perceived rigor, even if they are individually classified as minor.
Finally, interpret a rubric score around 3.5 not as a verdict but as a staging post. It reflects a draft that has a credible idea and basic structure, yet still needs one or two focused revision cycles. Authors who treat this feedback as an opportunity to stress-test their methods and clarify their claims are more likely to avoid surprises in formal peer review.
Want this level of scrutiny on your own manuscript?
ManuscriptMind runs the same review on your draft before you submit it. Five full reviews free, no credit card.
Get a free review