Fundamentals of search engines: indexing and ranking
Be able to describe indexing and ranking at a basic level — the first step towards retrieval.
Prerequisites
Everyday explanation
How can a search engine find among billions of pages in a tenth of a second?
The answer: it does not search when you search. It has already searched in advance.
Three things happen, and only the last one when you press search:
| Step | When | What |
|---|---|---|
| Crawling | constantly, in advance | robots follow links and fetch pages |
| Indexing | in advance | for each word, it notes which pages contain it |
| Ranking | when you search | the pages containing the words are sorted by how good they are |
The index is like the index at the back of a book. Instead of flipping through the whole book, you look up the word and get the page numbers directly.
"cat" → page 12, 45, 88
"dog" → page 45, 90
"cat food" → page 45
If you search for «cat dog», the search engine looks in its index and sees directly that page 45 has both. It never needs to open any page.
Intuition
Ranking is the hard part. A thousand pages may contain your words. Which one should be on top?
Search engines weigh many signals together:
| Signal | The question it answers |
|---|---|
| Word occurrence | are the words in the title or far down in the text? |
| Rarity | «dog» says more than «and» |
| Inbound links | do many other pages link here? (PageRank) |
| Freshness | does the date matter for this specific query? |
| Clicks | do people usually click and stay? |
| Location and language | is the page relevant where you are? |
Rare words weigh more is a key idea. The word «and» is on every page and helps nothing. The word «pterosaur» is on few — if the search engine finds it, it has found something.
Two people get different results for the same search. This depends on location, language, previous searches, and which device they use. It is practical — but it also means you and your colleague can see different versions of reality without noticing it.
Source criticism becomes a search engine question: the top results are those the ranking algorithm liked best, not the truest. Someone has also actively tried to get there. And ads are often at the top, marked but easy to miss.
Interactive
Build an index on paper. Take four short texts:
1: The cat sleeps on the mat
2: The dog eats food
3: The cat and the dog play
4: The mat is new
Step 1 — write the index. Go through word by word:
cat → 1, 3
sleeps → 1
mat → 1, 4
dog → 2, 3
eats → 2
food → 2
and → 3
play → 3
new → 4
is → 4
Step 2 — search. «cat dog» → text 1 and 3 for «cat», 2 and 3 for «dog». Both are in 3. It wins.
Step 3 — rank. If you only searched for «cat», you get 1 and 3. Which should be first? Now you need a rule. Suggestion: the one where the word makes up the largest share of the text.
Step 4 — find the problem. Search for «food». You get text 2. But text 2 is about the dog eating — and the word «food» is actually also inside «mat» in text 1 and 4, but as part of another word.
And vice versa: search for «cat food» and you get nothing, even though texts 1 and 2 together are about exactly that.
This is the Swedish compound word problem, and it is one of the things a modern search engine must solve — either by splitting compound words, or by comparing meaning instead of words. The latter is what AI search does, and it is the next step.
Mastery means
- Explains what an index is
- Describes how results are ranked
- Understands why search results differ between users
Sign in to do the exercises and build your mastery up.
Sources
- Internetstiftelsen — Internetkunskap — free to read
- CS Unplugged (CC BY-SA 4.0) — CC BY-SA 4.0
- Skolverket — About AI in school (in Swedish) — Skolverket's open terms