Visualizers

Learn by moving things

Interactive explainers for the ideas behind retrieval, embeddings, and agents. A formula tells you what is true; dragging a vector shows you why. Each one also appears inside the article it belongs to, and runs entirely in your browser, so drag, poke, and break it.

Portrait of Sachin Gupta rendered in binary

Inverted index & boolean retrieval

From: How Search Works

The data structure under almost every search engine: term → the documents that contain it. Toggle query terms and switch AND/OR to watch it resolve a query into a candidate set.

Query — click terms
Inverted index → postings
  • foxd0d2d4
  • fastd0d1d3
Documents1 match
  • d0fast red fox
  • d1red fast car
  • d2quiet fox den
  • d3red fast engine
  • d4blue fox tail
AND keeps documents in every selected term's list (intersection).

The default text-ranking function. Tune term frequency, document length, term rarity, and the k1/b knobs, and watch the score and its saturation curve respond.

BM25 term score
3.61
IDF 2.30 · ceiling 5.06
Score vs term frequency

Dashed line = the ceiling more frequency can never beat.

More of a term helps with diminishing returns (k1). Longer documents are penalized (b). Rarer terms (low rarity count) score higher (IDF).

How a search engine measures a typo. Type two words and the Levenshtein dynamic-programming matrix fills in live, with one cheapest edit path lit up and the edit distance in the corner.

Edit distance
3
ε
s
i
t
t
i
n
g
ε
0
1
2
3
4
5
6
7
k
1
1
2
3
4
5
6
7
i
2
2
1
2
3
4
5
6
t
3
3
2
1
2
3
4
5
t
4
4
3
2
1
2
3
4
e
5
5
4
3
2
2
3
4
n
6
6
5
4
3
3
2
3
Each cell = cheapest edits to match the two prefixes. Down = delete, right = insert, diagonal = substitute (free when the letters match). The lit path is one cheapest way; the bottom-right cell is the answer.
Visualizers — Sachin Gupta