How we measured AI writing across arXiv, and where the measurement breaks(unslop.run)
233 points by dopamine_daddy 23 hours ago | 159 comments
tl;dr: Researchers scored 12,750 arXiv papers with a detector calibrated so pre-ChatGPT papers flag at 0.4%, and found the machine-written share rose from that floor to ~32% recently, peaking near 39% in early 2026. Computer science leads at 65%, while math sits near 0.7%—though the authors caution that math's sparse prose may be out-of-distribution for the detector. Reported figures are a lower bound: detector coverage is imperfect, and flags capture "machine-like writing" (including heavy AI editing), not authorship.
HN Discussion:
  • Personal tests show pre-LLM writing flagged as AI, suggesting detector unreliability
  • Author restating their own methodology and findings
  • AI-assisted writing is genuinely widespread in academia and industry, confirming the trend
  • Text-only AI detection is fundamentally impossible or methodologically flawed
  • Detector may be picking up on post-2022 vocabulary/jargon shifts rather than actual AI writing