arXiv updated its rate limit policy on October 1, 2026. The change was first discussed in a Hacker News thread that reached 62 points and 19 comments.
What Changed in the Policy
arXiv introduced stricter request caps for bulk downloads and automated access. The policy targets scripts that previously pulled large volumes of PDFs and metadata without rate controls.
Researchers using manual browsing face no visible change. Automated tools and institutional scrapers now hit limits faster than before.
How the Limits Work
The update applies per-IP and per-user-agent rules. Exact numeric thresholds are not published in the announcement, but early reports indicate daily bulk requests are now restricted more tightly than the prior soft limits.
Community members testing the new rules report that simple wget or curl loops trigger blocks within minutes when run without delays.
Impact on Common Workflows
- Bulk metadata harvesting for training datasets now requires explicit delays or distributed IPs
- Institutional mirrors and search engines must register for higher quotas
- Individual researchers downloading under 100 papers per day remain unaffected
Who Should Adjust Their Setup
Teams building large-scale paper corpora or citation graphs need to add exponential backoff and respect the new caps. Solo users querying fewer than a few dozen papers daily can continue without changes.
Skip this update if your workflow is entirely manual or uses the official API with registered keys.
Comparison with Prior Limits
| Aspect | Old Policy | New Policy (Oct 2026) |
|---|---|---|
| Bulk daily requests | Soft, rarely enforced | Hard caps with blocks |
| Automated scripts | Minimal delays OK | Exponential backoff required |
| API registration | Optional | Recommended for high volume |
Practical Next Steps
Register for an API key on the arXiv site if your volume exceeds casual use. Add sleep intervals of at least 3 seconds between requests in any script. Monitor response headers for remaining quota signals.
Bottom line: The October 2026 update forces automated users to slow down or register, while leaving normal research access unchanged.
Top comments (0)