Back to news
Large Language Models
Jul 20, 2026

Study analyzes AI-generated content in arXiv papers, revealing significant prevalence

Jul 20, 2026
AI Summary

A study of 12,750 arXiv papers indicates that approximately 32% of new submissions are perceived as machine-written, with a notable increase following the introduction of ChatGPT. The research highlights the limitations of AI detection methods and the variability across different academic fields.

  • The study analyzed 12,750 papers from arXiv, focusing on submissions from 2021 to 2026.
  • A detector calibrated for academic writing was used, achieving a 0.4% false-positive rate for genuine human-written text.
  • The flagged share of machine-written text remained stable at 0.4% in 2021 and 2022, then increased to about 32% in 2026, peaking at 39% early in the year.
  • Computer science papers showed the highest prevalence of machine-like writing at approximately 65%, while mathematics papers had the lowest at around 0.7%.
  • The study's methodology involved scoring the full text of papers rather than just abstracts, as abstracts tend to underrepresent the machine-written signal.
  • Limitations include the coarse nature of per-field control rates and the potential for low scores in mathematics to reflect either low adoption of AI or detector insensitivity.
  • The detector cannot definitively determine authorship; it only assesses whether text appears machine-generated, which may include heavily AI-assisted editing.
  • The study emphasizes that the reported prevalence of machine-like writing is a lower bound, as the detector may not cover all AI generation methods effectively.
ai writingarxivmeasurementlanguage modelsresearch