返回首页
原创
原创观点
2026/10/08

The Cost of Infinite Content: Why arXiv is Hitting the Brakes

The scientific community has long relied on speed. Before peer review takes months or even years to validate a breakthrough, researchers share their early...

The Cost of Infinite Content: Why arXiv is Hitting the Brakes
学术出版
生成式AI
内容审核
arXiv
信息过载

The scientific community has long relied on speed. Before peer review takes months or even years to validate a breakthrough, researchers share their early findings on preprint servers like arXiv, accelerating the global pace of discovery. But what happens when the machinery of publication moves too fast, fueled not by human genius, but by artificial intelligence?

Recently, arXiv—the legendary open-access repository for physics, mathematics, and computer science—was forced to slow down. The platform announced strict new rate limits for authors, capping submissions at two per calendar month and allowing a maximum of three active submissions at any given time.

The culprit is a staggering influx of AI-generated academic spam, often referred to as "slop." According to Thomas Dietterich, an Oregon State University professor emeritus and chair of arXiv’s editorial advisory council, a relatively small fraction of authors are weaponizing AI to submit a massive volume of low-quality work. This flood is overwhelming the platform's volunteer moderation team. As a result, legitimate, high-quality research is being delayed by days or even weeks while human reviewers dig through the synthetic noise.

The numbers paint a stark picture of a system under siege. Over the past two years, total submissions to arXiv have doubled, with the computer science category alone experiencing a sixfold increase. Much of this surge is driven by authors using large language models to generate "salami papers"—a practice where a single, modest study is sliced into multiple thin, narrow articles to artificially inflate a researcher's publication record.

It is important to note that arXiv is not inherently anti-AI. The platform’s official policy permits the use of artificial intelligence as a research assistant—whether for parsing data, writing code, or drafting text—provided the use is disclosed and the final paper actually advances the field. The problem is that a vast number of recent submissions fail to clear this basic hurdle of scholarly value.

This new rate limit is just the latest defensive maneuver in an ongoing battle. Over the past year, arXiv has taken several unprecedented steps to protect its archives. It halted the acceptance of review and position papers in the computer science category, citing low-effort, LLM-written submissions. It also tightened the rules for new authors, requiring endorsements from established researchers rather than simply accepting an academic email address, and threatened one-year bans for those caught submitting AI spam.

The bottleneck at arXiv serves as a powerful microcosm for the broader internet. When the cost of generating convincing text drops to near zero, the burden of quality control shifts entirely to human moderators. The challenge moving forward isn't just about detecting AI-written text; it's about preserving the signal in a world increasingly drowned out by automated noise.

Key Points

  • arXiv has limited researchers to two submissions per month to combat a surge in low-quality papers.
  • A small number of authors using AI to generate 'salami papers' are overwhelming volunteer moderators.
  • Submissions have doubled in two years, causing delays for legitimate, high-quality research.
  • arXiv allows AI use if disclosed and valuable, but is aggressively fighting automated academic spam.

Why It Matters

The struggle to filter AI-generated academic spam highlights a broader societal challenge: as AI makes content creation effortless, the human cost of verifying and curating information is skyrocketing.


Sources: