Skip to content
AI-grafen
EUniversityData handling· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Web scraping — the technique and the rules

Be able to fetch data from the web responsibly: robots.txt, rate, terms.

Prerequisites

Intuition

The order to try, always:

  1. Is there an API? Use it. More stable, faster, allowed.
  2. Is there a downloadable dump? Wikipedia, Kolada, Statistics Sweden and many public authorities offer one.
  3. Only after that: scrape.

Before you scrape, check four things:

  • robots.txt — what may be fetched and how often.
  • The terms of use — they can forbid automated fetching even when robots.txt allows it.
  • The copyright in the content — being allowed to fetch is not being allowed to publish.
  • Personal data — data being public does not make it free to collect at scale.

Code

import time, urllib.robotparser as rp
import httpx
from bs4 import BeautifulSoup

UA = "AI-grafen-bot/1.0 (+https://example.org/about-the-bot; [email protected])"

class Scraper:
    def __init__(self, base, delay=2.0):
        self.base, self.delay, self.last = base, delay, 0.0
        self.robots = rp.RobotFileParser()
        self.robots.set_url(base.rstrip("/") + "/robots.txt")
        self.robots.read()
        self.client = httpx.Client(headers={"User-Agent": UA}, timeout=20, follow_redirects=True)

    def fetch(self, url):
        if not self.robots.can_fetch(UA, url):
            raise PermissionError(f"robots.txt does not allow {url}")
        wait = self.delay - (time.time() - self.last)
        if wait > 0:
            time.sleep(wait)                        # one page every two seconds
        r = self.client.get(url)
        self.last = time.time()
        if r.status_code == 429:
            time.sleep(60); return self.fetch(url)  # respect the backpressure
        r.raise_for_status()
        return r.text

s = Scraper("https://example.org")
soup = BeautifulSoup(s.fetch("https://example.org/articles"), "html.parser")

Six rules that separate responsible fetching from abuse:

  1. Identify yourself in the User-Agent, with contact details.
  2. One request at a time, with a pause. Scraping the same domain in parallel is in practice a denial-of-service attack.
  3. Respect 429 and Retry-After.
  4. Cache — never fetch the same page twice unnecessarily.
  5. Fetch only what you need.
  6. Save the source, the timestamp and the licence for every document — otherwise the material cannot be used responsibly afterwards.

The last one is exactly what AI-grafen's source register does: every source has a licence, a fingerprint, a fetch date and a responsible editor.

Mastery means

  • Fetches data from the web technically correctly
  • Follows robots.txt, the terms and a reasonable rate
  • Knows what is allowed and what is not

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences