Web scraping — the technique and the rules
Be able to fetch data from the web responsibly: robots.txt, rate, terms.
Prerequisites
- DAPIs and HTTPrequired
- DLicences and open datarequired
- DRegular expressionsrequired
Intuition
The order to try, always:
- Is there an API? Use it. More stable, faster, allowed.
- Is there a downloadable dump? Wikipedia, Kolada, Statistics Sweden and many public authorities offer one.
- Only after that: scrape.
Before you scrape, check four things:
robots.txt— what may be fetched and how often.- The terms of use — they can forbid automated fetching even when robots.txt allows it.
- The copyright in the content — being allowed to fetch is not being allowed to publish.
- Personal data — data being public does not make it free to collect at scale.
Code
import time, urllib.robotparser as rp
import httpx
from bs4 import BeautifulSoup
UA = "AI-grafen-bot/1.0 (+https://example.org/about-the-bot; [email protected])"
class Scraper:
def __init__(self, base, delay=2.0):
self.base, self.delay, self.last = base, delay, 0.0
self.robots = rp.RobotFileParser()
self.robots.set_url(base.rstrip("/") + "/robots.txt")
self.robots.read()
self.client = httpx.Client(headers={"User-Agent": UA}, timeout=20, follow_redirects=True)
def fetch(self, url):
if not self.robots.can_fetch(UA, url):
raise PermissionError(f"robots.txt does not allow {url}")
wait = self.delay - (time.time() - self.last)
if wait > 0:
time.sleep(wait) # one page every two seconds
r = self.client.get(url)
self.last = time.time()
if r.status_code == 429:
time.sleep(60); return self.fetch(url) # respect the backpressure
r.raise_for_status()
return r.text
s = Scraper("https://example.org")
soup = BeautifulSoup(s.fetch("https://example.org/articles"), "html.parser")
Six rules that separate responsible fetching from abuse:
- Identify yourself in the User-Agent, with contact details.
- One request at a time, with a pause. Scraping the same domain in parallel is in practice a denial-of-service attack.
- Respect 429 and Retry-After.
- Cache — never fetch the same page twice unnecessarily.
- Fetch only what you need.
- Save the source, the timestamp and the licence for every document — otherwise the material cannot be used responsibly afterwards.
The last one is exactly what AI-grafen's source register does: every source has a licence, a fingerprint, a fetch date and a responsible editor.
Mastery means
- Fetches data from the web technically correctly
- Follows robots.txt, the terms and a reasonable rate
- Knows what is allowed and what is not
Sign in to do the exercises and build your mastery up.
Sources
- Creative Commons — licenser — CC BY 4.0
- IMY — Integritetsskyddsmyndigheten — myndighetsmaterial